METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Sakana AI Releases Benchmark for AI Paper Review

Sakana AI has released a paper presenting a benchmark that measures how well AI review systems catch contradictions inside research papers, along with MLR, a review-assistance system. MLR caught 73.43% of contradictions that overturn a paper's core claim, while existing systems stopped at 11 to 15%.

Sakana AI Releases Benchmark for AI Paper Review

Image: METAL

Summary

  • On October 9, Sakana AI released "Beyond Imitation," a paper accepted for publication in TMLR. It pairs a benchmark that plants 1,164 contradictions into 257 papers with MLR, a Claude-based review-assistance system.
  • Under a four-review setting MLR caught 73.43% of contradictions that overturn a core claim, but simply switching LLM-Review's model to Claude raised its distance-0 detection rate from 14.56% to 35.40%, meaning much of the gap came from the model.
  • On real retracted papers the detection rate fell to 26.07%, one review cost about $0.50, and researchers agreed with 81% of comments across 38 sessions.

Sakana AI on October 9 released "Beyond Imitation," a paper on how AI can assist the people who review research papers. The paper introduces a benchmark that measures how well AI review systems catch contradictions inside papers, together with Multi-Layered Review (MLR), a review-assistance system Sakana AI designed. MLR caught 73.43% of contradictions that directly conflict with a paper's core claim, while the three existing AI review systems used for comparison stopped at 11 to 15%. The paper has been accepted for publication in the machine learning journal TMLR.

The starting point is review volume. According to a graphic Sakana AI released with the paper, submissions to the ICLR conference grew from 2,594 papers in 2020 to 19,525 in 2026, and abstract registrations for 2027 are estimated at about 62,000. At ICLR 2026, 18,054 reviewers wrote 76,139 reviews. On X, Sakana AI said "peer review needs support, not substitutes," adding that "much of their development and evaluation focuses on how closely they imitate human reviews." That is where the paper's title, "Beyond Imitation," comes from.

The benchmark was built from deliberately broken papers. The researchers gathered 257 papers published at ACL, AISTATS, CVPR and ICML in 2025 and at NeurIPS in 2024. They selected only papers under Creative Commons licenses that allow modification and redistribution. Using Gemini 2.5 Pro, they drew a knowledge graph for each paper linking claims, evidence, methodology and implementation, then picked one point at each distance from the core claim and planted a sentence that conflicts with the original. A distance of 0 means the core claim itself was overturned, the most severe kind of error. This produced 1,164 contradictions. OpenAI's o3 graded the results, averaged over 10 runs; it correctly answered "no" on unmodified original papers 99.9% of the time and recognized 86.8% of contradictions that had actually been caught.

사카나AI가 공개한 ICLR 제출 논문 추이 그래프. 2020년 2594편에서 2026년 1만9525편으로 늘고 2027년 초록 등록은 약 6만2000건으로 추정되며, ICLR 2026 리뷰 7만6139건과 심사위원 1만8054명이 표시돼 있다

MLR runs on three agents. An Appendix Agent (Claude Haiku 3.5) summarizes the appendix, and a Literature Review Agent (Claude Sonnet 4) organizes prior work through web search. A Review Agent (Claude Sonnet 4) handles the main text, up to 10 pages, and reads it three times. The first pass produces an outline, the second examines evidence and weaknesses, and the third writes the review by combining the earlier results. The design adapts Keshav's 2007 method of reading research papers in three passes. Sakana AI explained that "the idea is simple: understand the paper before judging it."

사카나AI가 공개한 벤치마크 구성 도식. 논문별 지식 그래프에서 핵심 주장과 이어진 근거 노드에 오류를 심고, 심사를 돌린 뒤 오류를 잡았는지 확인하는 세 단계가 그려져 있다

The results split at distance 0. Under a setting where four reviews are combined and detection counts as successful if any one of them flags the error, MLR caught 73.43% of distance-0 contradictions and 40.95% overall. The comparison systems scored 14.56% for LLM-Review, 11.17% for AI Reviewer and 14.81% for AgentReview. With a single review, MLR still recorded 60.79% at distance 0 and 28.32% overall. In every system, detection fell the further a contradiction sat from the core claim, and the researchers said this shows the severity scale based on knowledge-graph distance works as intended.

A large share of this gap came from the model. The three comparison systems all use OpenAI GPT-family models, while only MLR uses Claude. When the researchers swapped only LLM-Review's model for Claude Sonnet 4, its single-review detection rate jumped from 6.39% to 16.43% overall and from 14.56% to 35.40% at distance 0. A variant of MLR that merges the three passes into a single prompt also scored 24.81% on a 497-paper subset, not far from the original single-review 28.32%. In their OpenReview response, the researchers acknowledged that the three-pass structure had no substantial impact on either error detection or score correlation. The difference did widen, however, on hard-to-find errors at distance 3 and beyond. Action editor Quanming Yao, who handled the review, recommended "accept with minor revision" on July 8 and wrote that "the benchmark and verification-centric framing are timely and useful."

The numbers fell on real retracted papers. On WithdrarXiv-Check, a collection of arXiv papers withdrawn because of errors, MLR raised a problem similar to the retraction reason 26.07% of the time and matched it exactly 16.11% of the time. The runners-up were AgentReview at 18.48% and AI Reviewer at 9.00%, respectively. Limited to papers whose retraction reason mentions a "proof," MLR fell to 15.2% and 8.7%. In an audit by human raters, 34% of 50 planted contradictions were judged not to flow naturally from sentence to sentence, and 8% were judged to sound AI-generated. The researchers responded that "it makes sense that due to the synthetic nature of the contradictions, the errors introduced are easier to detect than genuine mistakes." They also offered the reading that existing systems miss even these easier errors.

They also measured distance from human review. On ICLR 2025 papers, the Pearson correlation between scores predicted from MLR reviews and actual human scores was 0.586, compared with 0.451 for NeurIPS 2024 and 0.429 for ICML 2025. At ICML, AI Reviewer edged ahead at 0.439. What the reviews focused on, however, differed from humans. Where human reviewers often pointed to clarity of writing and novelty, MLR put weight on validity, and its difference in focus distribution from human reviews (average KL divergence) was 0.240, the second largest among the systems compared after AI Reviewer's 0.266.

The system was also tested against attacks. The researchers inserted hidden text designed to sway AI reviewers after the conclusion of 50 rejected papers, most of which had been submitted to ICLR 2025. Every system was heavily swayed by the text, and MLR noticed the manipulative text in 8 of the 50 cases. Cost came to about $0.50 per review. MLR used 189,062 input tokens, about half of AI Reviewer's 403,654. The researchers estimated the API cost of all experiments in the paper at about $3,500.

Real researchers were hardest on the weaknesses section. Across 38 sessions in which researchers uploaded their own manuscripts to the full MLR system, participants agreed with 305 of 378 review comments (81%). Agreement was 94% for the overall recommendation and 92% for strengths, but lowest for weaknesses at 68%. The criticality score was 2.88 out of 5, close to neutral, and a weak negative correlation (r=-0.29) emerged in which the harsher the critique, the lower the helpfulness rating.

According to the 58-page paper and the OpenReview review record checked by METAL, the authors are five Sakana AI researchers: Rachel Teo, Yutaro Yamada, Shashank Kotyan, Yuki Imajuku and Tarin Clanuwat. The paper was posted to arXiv on October 8, and all three reviewers asked the authors to separate the effects of the model and the design, a result that went into a figure in the main text. METAL has previously reported that Sakana AI released PC-ALM, a learning method without backpropagation, and that it brought on Jürgen Schmidhuber as chief scientific advisor.

Seen through an engineer's eyes, what this paper leaves behind is the benchmark more than MLR. Much of the rise in detection came from swapping the model, and the effect of the three-pass structure was confined to hard errors. The pipeline that weighs and plants errors using knowledge graphs, by contrast, can keep expanding with new conference papers. Sakana AI said its goal is "to give reviewers useful support in checking research, with human expertise and judgment at the center."

Comments