paper-with-me

Papers

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

2026-05-28 · Krzysztof Żurawicki, Julia Farganus, Arkadiusz Gaweł, Mateusz Bystroński, Tomasz Jan Kajdanowicz arxiv

The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review process, particularly in terms of improving its speed and scalability. Yet, it remains unknown whether LLMs engage with scientific manuscripts in the same manner as human reviewers, or whether they merely produce review-looking text. To address this, we introduce the Peer Review AI Benchmark (PRAIB), a novel framework comprising thoroughly defined metrics that measure review specificity, style, and behavior of engagement. To complement the PRAIB framework, we conduct a large-scale empirical study leveraging a dataset of 11,000 reviews generated by five proprietary and open-source models for 1,000 ICLR and NeurIPS papers. Spanning the 2021--2025 period, these machine-generated reviews are compared against original human feedback across diverse prompting strategies to identify systematic behavioral divergences. Our analysis reveals that the generated reviews diverge significantly from feedback provided by human reviewers: LLM ratings are less variable, positively biased, and overconfident, and their cross-reference patterns are model-dependent and distinct from human norms. Furthermore, when evaluated through PRAIB, we observe that LLMs tend to generate longer, more complex reviews, yet frequently overlook the atomic weaknesses noted by human reviewers. By characterizing where and how LLMs reviewing behavior departs from human norms, PRAIB provides the community with a diagnostic tool for identifying which aspects of the review process LLMs can reliably support today and which require further development before deployment.

📄 PDF Abstract BibTeX arXiv:2605.29815

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

2026-04-15 · Joydeep Biswas, Sheila Schoepp, Gautham Vasan, Anthony Opipari 외 arxiv

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider …

Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review

2026-01-28 · Vibhhu Sharma, Thorsten Joachims, Sarah Dean arxiv

There are increasing indications that LLMs are not only used for producing scientific papers, but also as part of the peer review process. In this work, we provide the first comprehensive analysis of LLM use across the p…

Position on LLM-Assisted Peer Review: Addressing Reviewer Gap through Mentoring and Feedback

2026-01-14 · JungMin Yun, JuneHyoung Kwon, MiHyeon Kim, YoungBin Kim arxiv

The rapid expansion of AI research has intensified the Reviewer Gap, threatening the peer-review sustainability and perpetuating a cycle of low-quality evaluations. This position paper critiques existing LLM approaches t…

What Can Natural Language Processing Do for Peer Review?

2024-05-10 · Ilia Kuznetsov, Osama Mohammed Afzal, Koen Dercksen, Nils Dycke 외

The number of scientific articles produced every year is growing rapidly. Providing quality control over them is crucial for scientists and, ultimately, for the public good. In modern science, this process is largely del…

Articles

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

2026-08-04 · Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kryštof Olík 외 arxiv

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capa…