paper-with-me

홈 › Papers

How Closely Do LLM Reviews Align with Human Peer Review?

2026-08-04 · Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez arxiv

Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.

📄 PDF Abstract BibTeX arXiv:2608.03659

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pre-review to Peer review: Pitfalls of Automating Reviews using Large Language Models

2025-12-14 · Akhil Pandey Akella, Harish Varma Siravuri, Shaurya Rohatgi arxiv

Large Language Models are versatile general-task solvers, and their capabilities can truly assist people with scholarly peer review as \textit{pre-review} agents, if not as fully autonomous \textit{peer-review} agents. W…

PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality

2026-06-18 · Zeyuan Chen, Ziqing Yang, Yihan Ma, Michael Backes 외 arxiv

As academic submissions grow, the traditional peer review process struggles to keep up, raising concerns about quality and fairness. A trend of using large language models (LLMs) for assistance has emerged. In this work,…

Prompt Engineering

OpenReviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews

2024-12-16 · Maximilian Idahl, Zahra Ahmadi

We present OpenReviewer, an open-source system for generating high-quality peer reviews of machine learning and AI conference papers. At its core is Llama-OpenReviewer-8B, an 8B parameter language model specifically fine…

Language ModelingLanguage ModellingLarge Language Model

Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions

2024-06-09 · Cheng Tan, Dongxin Lyu, Siyuan Li, Zhangyang Gao 외

Large Language Models (LLMs) have demonstrated wide-ranging applications across various fields and have shown significant potential in the academic peer-review process. However, existing applications are primarily limite…

Review Generation

Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future

2026-04-30 · Sihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu 외 arxiv

Peer review is a multi-stage process involving reviews, rebuttals, meta-reviews, final decisions, and subsequent manuscript revisions. Recent advances in large language models (LLMs) have motivated methods that assist or…