paper-with-me

Papers

LLM-REVal: Can We Trust LLM Reviewers Yet?

2025-10-14 · Rui Li, Jia-Chen Gu, Po-Nien Kung, Heming Xia, Junfeng liu, Xiangwen Kong, Zhifang Sui, Nanyun Peng arxiv

The rapid advancement of large language models (LLMs) has inspired researchers to integrate them extensively into the academic workflow, potentially reshaping how research is practiced and reviewed. While previous studies highlight the potential of LLMs in supporting research and peer review, their dual roles in the academic workflow and the complex interplay between research and review bring new risks that remain largely underexplored. In this study, we focus on how the deep integration of LLMs into both peer-review and research processes may influence scholarly fairness, examining the potential risks of using LLMs as reviewers by simulation. This simulation incorporates a research agent, which generates papers and revises, alongside a review agent, which assesses the submissions. Based on the simulation results, we conduct human annotations and identify pronounced misalignment between LLM-based reviews and human judgments: (1) LLM reviewers systematically inflate scores for LLM-authored papers, assigning them markedly higher scores than human-authored ones; (2) LLM reviewers persistently underrate human-authored papers with critical statements (e.g., risk, fairness), even after multiple revisions. Our analysis reveals that these stem from two primary biases in LLM reviewers: a linguistic feature bias favoring LLM-generated writing styles, and an aversion toward critical statements. These results highlight the risks and equity concerns posed to human authors and academic research if LLMs are deployed in the peer review cycle without adequate caution. On the other hand, revisions guided by LLM reviews yield quality gains in both LLM-based and human evaluations, illustrating the potential of the LLMs-as-reviewers for early-stage researchers and enhancing low-quality papers.

📄 PDF Abstract BibTeX arXiv:2510.12367

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are Your Reviewers Being Treated Equally? Discovering Subgroup Structures to Improve Fairness in Spam Detection

2022-04-24 · Jiaxin Liu, Yuefei Lyu, Xi Zhang, Sihong Xie

User-generated reviews of products are vital assets of online commerce, such as Amazon and Yelp, while fake reviews are prevalent to mislead customers. GNN is the state-of-the-art method that detects suspicious reviewers…

FairnessSpam detection

Should we trust web-scraped data?

2023-08-04 · Jens Foerderer

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated compute…

What happens when reviewers receive AI feedback in their reviews?

2026-02-14 · Shiping Chen, Shu Zhong, Duncan P. Brumby, Anna L. Cox arxiv

AI is reshaping academic research, yet its role in peer review remains polarising and contentious. Advocates see its potential to reduce reviewer burden and improve quality, while critics warn of risks to fairness, accou…

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

2026-09-17 · Sy-Tuyen Ho, Minghui Liu, Furong Huang hf

Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, …

Trojan Horses in Amazon's Castle: Understanding the Incentivized Online Reviews

2018-06-06 · 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM) 2018 6 · S Jamshidi, R Rejaie, J. Li

During the past few years, sellers have increasingly offered discounted or free products to selected reviewers of e-commerce platforms in exchange for their reviews. Such incentivized (and often very positive) reviews ca…