paper-with-me

홈 › Papers

FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers

2025-11-26 · Sarina Xi, Vishisht Rao, Justin Payan, Nihar B. Shah arxiv

The identification and localization of errors is a core task in peer review, yet the exponential growth of scientific output has made it increasingly difficult for human reviewers to reliably detect errors given the limited pool of experts. Recent advances in Large Language Models (LLMs) have sparked interest in their potential to support such evaluation tasks, from academic peer review to automated scientific assessment. However, despite the growing use of LLMs in review systems, their capabilities to pinpoint errors remain underexplored. In this work, we introduce Fault Localization Across Writing in Science (FLAWS), an automated benchmark consisting of 713 paper-error pairs designed to evaluate how effectively LLMs detect errors that undermine key claims in research papers. We construct the benchmark by systematically inserting claim-invalidating errors into peer-reviewed papers using LLMs, paired with an automated evaluation metric that measures whether models can identify and localize these errors. Developing such a benchmark presents unique challenges that we overcome: ensuring that the inserted errors are well-defined, challenging, and relevant to the content of the paper, avoiding artifacts that would make identification trivial, and designing a scalable, automated evaluation metric. On the resulting benchmark, we evaluate five frontier LLMs: Claude Sonnet 4.5, DeepSeek Reasoner v3.1, Gemini 2.5 Pro, GPT 5, and Grok 4. Among these, GPT 5 is the top-performing model, achieving 39.1% identification accuracy when k=10, where k is the number of top-ranked error text candidates generated by the LLM.

📄 PDF Abstract BibTeX arXiv:2511.21843

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks

2026-02-05 · Nishant Balepur, Bhavya Rajasekaran, Jane Oh, Michael Xie 외 arxiv

Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-inspired toolkit using LLM judges to flag three common MCQ flaws: 1) contam…

Question Answering

Towards Automating Scientific Review with Google's Paper Assistant Tool

2026-06-26 · Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes 외 arxiv

Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challen…

Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

2026-04-15 · Zipeng Ling, Shuliang Liu, Shenghong Fu, Yuehao Tang 외 arxiv

LLM reasoning traces suffer from complex flaws -- *Step Internal Flaws* (logical errors, hallucinations, etc.) and *Step-wise Flaws* (overthinking, underthinking), which vary by sample. A natural approach would be to pro…

Mathematical Reasoning

Towards Characterizing Scientific Image Utility and Upgradability

2026-06-02 · WenZhe Li, Qihang Yan, Liang Chen, Junying Wang 외 arxiv

Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content that introduces subtle but consequential errors. Existing evaluation pa…

Automatic Vertebra Localization and Identification in CT by Spine Rectification and Anatomically-constrained Optimization

2020-12-14 · CVPR 2021 1 · Fakai Wang, Kang Zheng, Le Lu, Jing Xiao 외

Accurate vertebra localization and identification are required in many clinical applications of spine disorder diagnosis and surgery planning. However, significant challenges are posed in this task by highly varying path…