paper-with-me

홈 › Papers

Automatically Evaluating the Paper Reviewing Capability of Large Language Models

2025-02-24 · Hyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, Juho Kim

Peer review is essential for scientific progress, but it faces challenges such as reviewer shortages and growing workloads. Although Large Language Models (LLMs) show potential for providing assistance, research has reported significant limitations in the reviews they generate. While the insights are valuable, conducting the analysis is challenging due to the considerable time and effort required, especially given the rapid pace of LLM developments. To address the challenge, we developed an automatic evaluation pipeline to assess the LLMs' paper review capability by comparing them with expert-generated reviews. By constructing a dataset consisting of 676 OpenReview papers, we examined the agreement between LLMs and experts in their strength and weakness identifications. The results showed that LLMs lack balanced perspectives, significantly overlook novelty assessment when criticizing, and produce poor acceptance decisions. Our automated pipeline enables a scalable evaluation of LLMs' paper review capability over time.

📄 PDF Abstract BibTeX arXiv:2502.17086

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews

2024-08-19 · Keith Tyser, Ben Segev, Gaston Longhitano, Xin-Yu Zhang 외

Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human rev…

Ethics

A Survey on Multimodal Benchmarks: In the Era of Large AI Models

2024-09-21 · Lin Li, Guikun Chen, Hanrong Shi, Jun Xiao 외

The rapid evolution of Multimodal Large Language Models (MLLMs) has brought substantial advancements in artificial intelligence, significantly enhancing the capability to understand and generate multimodal content. While…

BenchmarkingSurvey

Gen-Review: A Large-scale Dataset of AI-Generated (and Human-written) Peer Reviews

2025-10-24 · Luca Demetrio, Giovanni Apruzzese, Kathrin Grosse, Pavel Laskov 외 arxiv

How does the progressive embracement of Large Language Models (LLMs) affect scientific peer reviewing? This multifaceted question is fundamental to the effectiveness -- as well as to the integrity -- of the scientific pr…

Deep Learning for Video Classification and Captioning

2016-09-22 · Zuxuan Wu, Ting Yao, Yanwei Fu, Yu-Gang Jiang

Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of today's big data. In this paper, we focus…

ClassificationDeep LearningGeneral ClassificationSentence+2

AutoMSC: Automatic Assignment of Mathematics Subject Classification Labels

2020-05-25 · Moritz Schubotz, Philipp Scharpf, Olaf Teschke, Andreas Kuehnemund 외

Authors of research papers in the fields of mathematics, and other math-heavy disciplines commonly employ the Mathematics Subject Classification (MSC) scheme to search for relevant literature. The MSC is a hierarchical a…

ArticlesClassificationGeneral ClassificationMath+1