10 분 소요

0. Introduction

Paper link

Towards Automating Scientific Review with Google’s Paper Assistant Tool은 peer review automation을 hype가 아니라 workflow design 문제로 다루는 논문이다. 논문이 제기하는 문제는 단순하다. AI가 과학 생산 속도를 높이고 있다면, 검증과 리뷰 process도 AI로 보조하지 않으면 병목이 더 커진다.

특히 ML과 CS conference의 submission volume은 빠르게 커지고 있다. Human reviewer가 line-by-line proof verification, experimental design audit, baseline comparison check, code/claim consistency check를 모두 처리하기에는 cognitive cost가 너무 크다. 하지만 그렇다고 AI에게 acceptance decision을 넘기는 것은 너무 위험하다. 이 논문은 그 사이 지점에 Paper Assistant Tool, 이하 PAT를 둔다.

PAT는 reviewer replacement가 아니라 author-facing pre-submission assistant다. Full manuscript를 ingest하고, segmenter가 paper를 semantic section으로 나누고, segment complexity에 따라 compute budget을 다르게 배정하고, Deep Review agents가 각 segment를 깊게 검증한 뒤, synthesis agent가 critique를 중복 제거하고 근거에 연결한다. 목표는 paper를 accept/reject하는 것이 아니라, 이론적 오류, 실험상 confound, missing comparison, proof gap, clarity issue를 authors가 submission 전에 고칠 수 있게 하는 것이다.

한 줄 요약: 이 논문은 scientific peer review bottleneck을 AI-assisted validation problem으로 보고, Paper Assistant Tool (PAT)을 segmenter, adaptive budgeting, deep review agents, global synthesis로 구성된 agentic review pipeline으로 제안하며, SPOT Math/CS proof-error subset과 STOC/ICML pre-submission pilot을 통해 author-facing review assistant의 가능성과 한계를 분석한다.

이 논문을 지금 볼 가치가 있는 이유는 다음과 같음.

  • AI-assisted science가 늘어날수록 generation보다 verification이 병목이 된다는 문제를 선명하게 제시한다.
  • AI review를 Role 1 author tool부터 Role 4 total automation까지 단계별 taxonomy로 정리한다.
  • PAT를 single model call이나 Pass@k가 아니라 segmented inference-scaling pipeline으로 설계한다.
  • SPOT benchmark Math/CS Equation/Proof subset에서 PAT가 zero-shot baseline보다 높은 recall을 보인다고 보고한다.
  • STOC와 ICML pilot에서 실제 author feedback을 수집해 pre-submission tool로서의 실용성을 보여준다.
  • Hallucinated critiques, PDF parsing, outdated knowledge, accountability, reviewer deskilling 같은 governance issue를 같이 다룬다.

이 글에서는 PAT 논문을 “AI가 review를 대체한다”가 아니라, scientific review system 안에서 AI agent가 어떤 role을 맡아야 안전하고 유용한가를 탐색한 workflow와 governance 논문으로 읽는다.

1. Problem Setting

1-1. Problem definition

Scientific peer review는 다음 일을 동시에 요구한다.

  • Theoretical claim을 확인한다.
  • Proof와 derivation을 검증한다.
  • Experimental design을 audit한다.
  • Baseline과 comparison을 확인한다.
  • 빠진 ablation을 찾는다.
  • Factual error나 citation error를 감지한다.
  • Clarity와 structure를 평가한다.
  • Constructive feedback을 제공한다.

AI-assisted generation이 늘어나면 manuscript volume과 complexity가 같이 증가한다. Review capacity는 그만큼 빨리 늘지 않는다. 따라서 validation bottleneck이 생긴다.

이를 간단히 쓰면 다음과 같다.

\[\mathrm{Science\ throughput} = \min \left( \mathrm{Generation\ capacity}, \mathrm{Validation\ capacity} \right)\]

Generation capacity가 validation capacity보다 빠르게 커지면 scientific process는 review-bound 상태가 된다.

PAT의 문제 설정은 AI에게 publication decision을 맡기는 것이 아니다. 더 좁다.

Submission 전에 authors에게 deep technical feedback을 제공해 obvious flaw와 non-obvious flaw를 early catch할 수 있는가?

1-2. Why previous approaches are insufficient

1) Single model call

Full paper를 한 번에 넣고 “review this paper”라고 하면 간단하다. 하지만 deep manuscript review에는 많은 thinking token이 필요하다. 긴 proof, dense equation, appendix, experiment, figure, code-like detail은 한 번의 call이 처리할 수 있는 effective reasoning capacity를 넘을 수 있다.

2) Naive Pass@k

여러 independent review는 recall을 높일 수 있다. 하지만 precision을 낮추고 deduplication burden을 만든다. 각 pass가 많은 possible issue를 만들면 author는 hallucinated critique와 redundant critique를 직접 걸러야 한다.

3) Human-only peer review

Human review는 여전히 필요하지만 submission surge에 맞춰 scale되기 어렵다. 특히 math-heavy CS paper에서는 line-by-line proof checking만으로도 며칠이 걸릴 수 있다.

4) Fully automated acceptance decision

Total automation은 언젠가 논의될 수 있지만, 현재 system은 critique를 hallucinate하거나 conceptual novelty를 놓치거나 scientific judgment를 지나치게 중앙화할 수 있다. 이 논문은 outcome에 대한 human control을 보존해야 한다고 주장한다.

2. Core Idea

2-1. Main contribution

논문의 기여는 다음과 같다.

  1. PAT pipeline
    • Segmenter
    • Adaptive compute budgeting
    • Deep Review agents
    • Global synthesis와 grounding
  2. SPOT case study
    • Math/CS Equation/Proof error subset을 평가한다.
    • Original SPOT SOTA, zero-shot Gemini, PAT를 비교한다.
  3. STOC와 ICML pilot programs
    • PAT를 author-facing pre-submission tool로 사용한다.
    • 두 conference에서 4,700개 이상의 submission을 review한다.
    • Author survey와 qualitative feedback을 수집한다.
  4. Four-role taxonomy for AI in peer review
    • Role 1: Tool for authors
    • Role 2: Tool for reviewers
    • Role 3: Supporting reviewer
    • Role 4: Total AI automation of peer review

2-2. Design intuition

PAT의 design은 단순한 관찰에서 출발한다.

Deep review는 하나의 prompt가 아니다. Attention을 관리해서 배분하는 과정이다.

Section마다 필요한 reasoning effort가 다르다.

  • Intro나 related work는 light review만으로 충분할 수 있다.
  • Theorem/proof section은 높은 thinking budget이 필요할 수 있다.
  • Experiment section은 baseline과 confound check가 필요하다.
  • Appendix는 targeted verification이 필요할 수 있다.

그래서 PAT는 manuscript를 semantic segment로 나누고 complexity에 따라 compute를 배정한다.

이는 단순한 효율화가 아니라 precision control이기도 하다. Segment-level review는 specialized analysis를 가능하게 하고, final synthesis는 critique를 중복 제거하고 근거에 연결한다.

3. Architecture / Method

3-1. Overview

Item Description
Goal Author-facing pre-submission scientific review assistant
Core system Paper Assistant Tool
Backbone Gemini Deep Think family와 proprietary inference-scaling pipeline
Stage 1 Document segmentation
Stage 2 Adaptive compute budgeting
Stage 3 Deep Review agents
Stage 4 Global synthesis와 grounding
Target domains Math, TCS, ML, CS papers
Evaluation SPOT proof-error subset, STOC/ICML pilots
Main role Role 1, author를 위한 AI tool

3-2. Module breakdown

1) Segmenter

Segmenter는 manuscript를 semantic component로 나눈다. Segment는 서로 overlap될 수도 있고, 문서 안에서 contiguous하지 않을 수도 있다. 예시는 다음과 같다.

  • Introduction
  • Related work
  • Theory
  • Proofs
  • Methodology
  • Experiments
  • Appendix

목표는 fixed size page chunking이 아니다. Review need에 맞춘 logical segmentation이다.

2) Adaptive budgeting

각 segment는 information density와 complexity에 따라 compute budget을 받는다.

Segment type Likely budget
Intro와 conclusion Light thinking
Method와 experiments Medium thinking
Theory와 proofs High thinking

이는 manuscript region마다 필요한 verification depth가 다르다는 문제를 해결한다.

3) Deep Review agents

Specialized review agent는 각 segment를 검사하되 full paper context에도 접근한다. Proof나 experiment issue는 manuscript 다른 곳의 definition이나 claim에 의존할 수 있기 때문이다.

Deep Review agent는 one-shot shallow reading이 아니라 inference scaling을 사용하도록 설계된다.

4) Global synthesis

Synthesis agent는 segment report를 합치고, duplicate를 제거하고, severity를 확인하며, search grounding을 사용해 존재하지 않는 paper나 theorem 같은 hallucinated critique를 줄인다.

최종 output은 binary accept/reject decision이 아니라 comprehensive review다.

3-3. Why segmentation beats naive scaling

Naive Pass@k는 같은 paper region을 여러 번 다시 보면서도 어려운 section은 under-review 상태로 남길 수 있다. PAT segmentation은 coverage를 강제하고 targeted budget allocation을 가능하게 한다.

개념적으로는 다음과 같다.

\[B_{\mathrm{total}} = \sum_{s \in S} B_s\]

여기서 $B_s$는 segment $s$에 배정된 compute budget이다. PAT는 budget을 uniform하게 뿌리지 않고 high-risk section에 더 많은 budget을 배정한다.

Scientific review에서는 error density가 paper 전체에 균일하지 않기 때문에, 이런 abstraction이 적절하다.

4. Training / Data / Recipe

4-1. Data and deployment context

논문은 두 가지 주요 empirical context를 보고한다.

Context Description
SPOT Math/CS subset 26 papers, 29 equation/proof errors
STOC와 ICML pilots 4,700개 이상의 submission을 PAT가 review

SPOT subset은 broader SPOT benchmark에서 Math/CS Equation/Proof error만 분리해 만든 것이다. Full SPOT benchmark에는 figure duplication 같은 multimodal error가 많이 포함되기 때문이다.

4-2. SPOT evaluation

논문은 다음을 비교한다.

Method Detection accuracy
Original SPOT SOTA 21.1%
Gemini 3.1 Pro zero-shot 55.2%
PAT with Gemini 3.1 Pro 89.7%

저자들은 자신들의 grading protocol이 specialized logic-aware grader와 human audit을 사용하므로, original SPOT paper의 strict keyword-matching grader와 수치를 직접 비교하기는 어렵다고 명시한다.

Zero-shot 55.2%에서 PAT 89.7%로 올라간 차이가 abstract에서 강조하는 34 percentage point recall improvement다.

4-3. Pilot programs

PAT는 두 pilot program에서 submission 전에 author에게 제공되었다.

Program Timing Role
STOC 2026 pilot November 2025 Math-heavy pre-submission tool
ICML pilot January 2026 Broader ML manuscript critique tool

두 deployment 모두 author-facing이었다. PAT는 formal acceptance decision의 일부가 아니었다.

Survey result는 다음을 포함한다.

Survey item STOC ICML
Would use PAT again 97% 92.1%
Improved clarity or readability 85.1% 87.0%
Education value 75.2% 83.9%
Very or mostly helpful 92.7% 90.7%
Feedback mostly or all grounded 55.8% 64.8%
Identified substantive theory gaps 11.6% 35.4%
Ran new experiments N/A 31%

이 수치는 independent peer-review outcome이 아니라 author survey response다. 즉 perceived usefulness와 actionability를 보여주는 지표로 읽어야 한다.

4-4. Engineering notes

  1. Logical role에 따라 segment한다
    • Fixed page chunk보다 theory, method, experiment를 구분하는 segment가 더 적절하다.
  2. Budget을 adaptive하게 배정한다
    • Proof와 technical section에는 더 높은 thinking budget을 쓴다.
  3. Synthesis로 precision을 제어한다
    • Critique를 deduplicate하고 claim을 search grounding으로 확인한다.
  4. Decision authority는 human에게 둔다
    • PAT는 Role 1 또는 Role 2 tool로 쓰는 것이 가장 안전하다.
  5. Hallucinated critique를 별도로 추적한다
    • 잘못된 critique는 reviewer와 author의 시간을 낭비하게 만들 수 있다.
  6. Grader caveat 없이 SPOT score를 비교하지 않는다
    • Logic-aware grading은 strict keyword matching과 다르다.

5. Evaluation

5-1. SPOT result

PAT는 filtered Math/CS Equation/Proof subset에서 89.7% detection accuracy를 달성한다. 비교 기준은 zero-shot Gemini 3.1 Pro의 55.2%, Original SPOT SOTA의 21.1%다.

이 결과는 specialized inference-scaling pipeline이 single call보다 깊은 proof error를 더 잘 찾을 수 있다는 논문의 main technical claim을 뒷받침한다.

다만 grading 방식이 original SPOT protocol과 다르다는 caveat가 중요하다.

5-2. STOC와 ICML author feedback

Pilot survey는 특히 “would use again”과 “very or mostly helpful” 항목에서 긍정적이다. Improved clarity/readability를 보고한 author 비율이 높다는 점은, PAT가 fatal error를 찾지 못하더라도 paper revision에 유용할 수 있음을 시사한다.

가장 중요한 practical finding은 다음과 같다.

  • STOC respondent의 11.6%가 substantive theory gap을 보고했다.
  • ICML respondent의 35.4%가 substantive theory gap을 보고했다.
  • ICML respondent의 31%는 PAT feedback 이후 새로운 experiment를 수행했다.

이들은 user-reported outcome이지만, PAT가 실제 paper revision behavior에 영향을 주었음을 보여준다.

5-3. Four-role taxonomy

논문은 네 가지 role을 제안한다.

Role Description Main risk
Role 1 Author를 위한 AI tool Paper가 polished해 보여도 deeper issue가 남을 수 있음
Role 2 Reviewer를 위한 AI tool Reviewer가 AI-generated critique를 검증해야 함
Role 3 Supporting reviewer 역할의 AI Hallucinated technical review가 decision에 영향을 줄 수 있음
Role 4 Total AI automation Centralized bias, deskilling, accountability issue

Pilot에서 PAT는 Role 1에 해당한다. 논문은 field가 Role 3이나 Role 4로 이동하기 전에 신중해야 한다고 주장한다.

5-4. What really matters in the experiments

1) 이 논문에서 PAT는 reviewer replacement가 아니다

가장 안전한 해석은 author를 위한 pre-submission validation assistant다.

2) Deep review에는 inference scaling이 중요하다

Single call은 proof-level issue를 놓칠 수 있다. Structured segmentation과 budget allocation은 coverage를 높인다.

3) Grounding과 synthesis가 필요하다

여러 review agent는 recall을 높이지만 precision을 낮출 수 있다. Synthesis와 search grounding이 중요하다.

4) Survey feedback은 ground truth가 아니다

Author feedback은 usefulness를 측정할 뿐, 모든 critique의 independent correctness를 보장하지 않는다.

6. Limitations

  1. 모든 detail이 reproducible한 것은 아니다
    • PAT는 proprietary Gemini Deep Think와 proprietary inference-scaling pipeline을 사용한다.
  2. SPOT subset이 작다
    • 26개 paper와 29개 error는 유용하지만 제한적이다.
  3. Grading protocol이 original SPOT과 다르다
    • Logic-aware grader와 human audit을 사용하므로 Original SPOT SOTA와 숫자를 직접 비교하기 어렵다.
  4. Author survey는 self-reported다
    • 이는 perceived usefulness와 revision behavior를 측정한 것이지, 독립적인 final paper quality를 측정한 것은 아니다.
  5. Hallucinated critique risk가 남아 있다
    • 논문은 proof error에 대한 false claim을 명시적인 limitation으로 언급한다.
  6. PDF parsing issue가 있다
    • Scientific paper에는 equation, figure, table, appendix, reference가 포함되어 있어 안정적으로 parse하기 어렵다.
  7. Knowledge freshness 문제
    • Date hallucination과 outdated knowledge cutoff가 challenge로 보고된다.
  8. Governance risk가 있다
    • Role 1에서 Role 3 또는 Role 4로 이동하면 accountability와 career에 대한 권한 구조가 바뀐다.
  9. Adversarial gaming이 가능하다
    • Author가 scientific truth보다 review agent를 만족시키도록 paper를 최적화할 수 있다.
  10. Compute equity 문제가 있다
    • 고품질 AI review에 큰 inference budget이 필요하다면 access inequality가 악화될 수 있다.

7. My Take

7-1. Why this matters for my work

PAT의 가장 중요한 메시지는 “AI가 peer review를 대체할 수 있다”가 아니다. 더 중요한 점은 scientific review도 agentic verification workflow로 설계되어야 한다는 것이다.

Paper review는 하나의 inference call이 아니다. Segmentation, targeted verification, critique grounding, deduplication, severity estimation, human decision support가 결합된 과정이다.

이 구조는 논문 리뷰뿐 아니라 technical report review, benchmark audit, model card review, safety evaluation, code review에도 그대로 적용될 수 있다.

7-2. Reuse potential

Research paper pre-check

Author는 submission 전에 PAT-like system을 사용해 proof gap, missing baseline, clarity issue를 찾을 수 있다. 가장 유용한 role은 reviewer replacement가 아니라 author-side debugging assistant다.

Internal research review

Lab은 internal paper approval 전에 AI pre-review를 돌릴 수 있다. 그러면 human reviewer는 first-pass error detection보다 novelty, framing, judgment에 시간을 더 쓸 수 있다.

Benchmark and evaluation audit

PAT pipeline은 benchmark auditor와 닮았다. Artifact를 segment하고, high-risk part에 compute를 배정하고, claim을 깊게 확인한 뒤 finding을 종합한다.

Education

논문은 student mentor로서의 AI 가능성도 시사한다. PAT-like system은 proof gap과 experimental design issue를 learning feedback으로 설명할 수 있다.

7-3. Production considerations

  • AI review라는 점을 명확히 label한다.
  • Objective error detection과 subjective acceptance recommendation을 분리한다.
  • False critique rate를 추적한다.
  • 각 critique에 citation, page reference, evidence를 제공한다.
  • Author가 hallucinated critique에 이의를 제기할 수 있게 한다.
  • Publication decision에 대한 human accountability를 유지한다.
  • AI review access가 inequality를 만드는지 audit한다.

7-4. Follow-up papers

  • SPOT benchmark
  • AI Scientist and AI Scientist-v2
  • Towards automating scientific discovery papers
  • LLM-as-a-Judge reliability studies
  • Automated theorem proving and proof checking
  • Peer review consistency experiments
  • Scientific claim verification and paper review assistant systems

8. Summary

  • PAT는 autonomous acceptance system이 아니라 agentic scientific review assistant다.
  • Manuscript를 segment하고, compute를 adaptive하게 배정하고, deep review agent를 실행한 뒤 grounded critique를 종합한다.
  • Filtered SPOT Math/CS proof error에서 PAT는 89.7% detection accuracy를 보고하며, zero-shot Gemini는 55.2%다.
  • STOC와 ICML pilot은 author가 PAT를 pre-submission revision에 유용하다고 느낄 수 있음을 보여준다.
  • 가장 중요한 open issue는 governance다. AI assistance가 어디서 끝나고 human responsibility가 어디서 시작되는지 정해야 한다.

댓글남기기