paper-with-me

Papers

PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality

2026-06-18 · Zeyuan Chen, Ziqing Yang, Yihan Ma, Michael Backes, Yang Zhang arxiv

As academic submissions grow, the traditional peer review process struggles to keep up, raising concerns about quality and fairness. A trend of using large language models (LLMs) for assistance has emerged. In this work, we take a critical step toward improving the quality of LLM-generated reviews. We propose the PeerCheck framework, which investigates LLM-human review differences (RQ1) and explores methods to improve LLM-generated review quality (RQ2). We first analyzed the human-written reviews with reviews generated by various LLMs and found that LLMs and humans focus on different terms, e.g., LLMs prioritize theory while humans emphasize methodology and experiments. We further adopt prompt engineering, such as Chain-of-Thought (CoT), and utilize retrieval-augmented generation (RAG) to enhance the LLM-generated reviews towards human-level quality. We find CoT significantly improves the quality of LLM reviews, while we discover an unexpected "RAG paradox," i.e., experiments with RAG produce different results for various LLMs and, in some cases, even reduce review quality. Our comprehensive analysis of LLM-generated academic reviews illustrates both possibilities and limitations, contributing to a more effective, human-aligned review system. Our dataset is available on https://github.com/TrustAIRLab/PeerCheck.

📄 PDF Abstract BibTeX arXiv:2606.20897

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

ReviewEval: An Evaluation Framework for AI-Generated Reviews

2025-02-17 · Chavvi Kirtani, Madhav Krishan Garg, Tejash Prasad, Tanmay Singhal 외

The escalating volume of academic research, coupled with a shortage of qualified reviewers, necessitates innovative approaches to peer review. While large language model (LLMs) offer potential for automating this process…

Language ModelingLanguage ModellingLarge Language Model

ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews

2025-03-11 · Xian Gao, Jiacheng Ruan, Jingsheng Gao, Ting Liu 외

Academic paper review is a critical yet time-consuming task within the research community. With the increasing volume of academic publications, automating the review process has become a significant challenge. The primar…

Comment Generation

Detecting AI-Generated Content in Academic Peer Reviews

2026-01-30 · Siyuan Shen, Kai Wang arxiv

The growing availability of large language models (LLMs) has raised questions about their role in academic peer review. This study examines the temporal emergence of AI-generated content in peer reviews by applying a det…

LLM-REVal: Can We Trust LLM Reviewers Yet?

2025-10-14 · Rui Li, Jia-Chen Gu, Po-Nien Kung, Heming Xia 외 arxiv

The rapid advancement of large language models (LLMs) has inspired researchers to integrate them extensively into the academic workflow, potentially reshaping how research is practiced and reviewed. While previous studie…

ReviewGuard: Enhancing Deficient Peer Review Detection via LLM-Driven Data Augmentation

2025-10-18 · Haoxuan Zhang, Ruochi Li, Sarthak Shrestha, Shree Harshini Mamidala 외 arxiv

Peer review serves as the gatekeeper of science, yet the surge in submissions and widespread adoption of large language models (LLMs) in scholarly evaluation present unprecedented challenges. While recent work has focuse…

Data AugmentationText Detection