paper-with-me

Papers

ReviewScore: Misinformed Peer Review Detection with Large Language Models

2025-09-25 · Hyun Ryu, Doohyuk Jang, Hyemin S. Lee, Joonhyun Jeong, Gyeongman Kim, Donghyeon Cho, Gyouk Chu, Minyeong Hwang, Hyeongwon Jang, Changhun Kim, Haechan Kim, Jina Kim, Joowon Kim, Yoonjeon Kim, Kwanhyung Lee, Chanjae Park, Heecheol Yun, Gregor Betz, Eunho Yang arxiv

Peer review serves as a backbone of academic research, but in most AI conferences, the review quality is degrading as the number of submissions explodes. To reliably detect low-quality reviews, we define misinformed review points as either "weaknesses" in a review that contain incorrect premises, or "questions" in a review that can be already answered by the paper. We verify that 15.2% of weaknesses and 26.4% of questions are misinformed and introduce ReviewScore indicating if a review point is misinformed. To evaluate the factuality of each premise of weaknesses, we propose an automated engine that reconstructs every explicit and implicit premise from a weakness. We build a human expert-annotated ReviewScore dataset to check the ability of LLMs to automate ReviewScore evaluation. Then, we measure human-model agreements on ReviewScore using eight current state-of-the-art LLMs. The models show F1 scores of 0.4--0.5 and kappa scores of 0.3--0.4, indicating moderate agreement but also suggesting that fully automating the evaluation remains challenging. A thorough disagreement analysis reveals that most errors are due to models' incorrect reasoning. We also prove that evaluating premise-level factuality shows significantly higher agreements than evaluating weakness-level factuality.

📄 PDF Abstract BibTeX arXiv:2509.21679

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Misinformation Propagation in Benign Multi-Agent Systems

2026-06-15 · Jonas Becker, Jan Philip Wahle, Terry Ruas, Bela Gipp arxiv

Multi-agent systems, in which multiple large language model agents solve problems through turn-based interaction, are increasingly deployed in high-stakes settings such as medical diagnosis, legal analysis, and forensic …

Medical Diagnosis

PeerPrism: Peer Evaluation Expertise vs Review-writing AI

2026-04-16 · Soroush Sadeghian, Alireza Daqiq, Radin Cheraghi, Sajad Ebrahimi 외 arxiv

Large Language Models (LLMs) are increasingly used in scientific peer review, assisting with drafting, rewriting, expansion, and refinement. However, existing peer-review LLM detection methods largely treat authorship as…

Text Detection

Is Your Paper Being Reviewed by an LLM? A New Benchmark Dataset and Approach for Detecting AI Text in Peer Review

2025-02-26 · Sungduk Yu, Man Luo, Avinash Madusu, Vasudev Lal 외

Peer review is a critical process for ensuring the integrity of published scientific research. Confidence in this process is predicated on the assumption that experts in the relevant domain give careful consideration to …

BenchmarkingText Detection

Detecting AI-Generated Content in Academic Peer Reviews

2026-01-30 · Siyuan Shen, Kai Wang arxiv

The growing availability of large language models (LLMs) has raised questions about their role in academic peer review. This study examines the temporal emergence of AI-generated content in peer reviews by applying a det…

CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection

2025-08-28 · Yihan Chen, Jiawei Chen, Guozhao Mo, Xuanang Chen 외 arxiv

The growing integration of large language models (LLMs) into the peer review process presents potential risks to the fairness and reliability of scholarly evaluation. While LLMs offer valuable assistance for reviewers wi…

Multi-Task Learning