paper-with-me

홈 › Papers

Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews

2026-04-21 · Bowen Li, Haochen Ma, Yuxin Wang, Jie Yang, Yining Zheng, Xinchi Chen, Xuanjing Huang, Xipeng Qiu arxiv

The rapid adoption of Large Language Models (LLMs) has spurred interest in automated peer review; however, progress is currently stifled by benchmarks that treat reviewing primarily as a rating prediction task. We argue that the utility of a review lies in its textual justification--its arguments, questions, and critique--rather than a scalar score. To address this, we introduce Beyond Rating, a holistic evaluation framework that assesses AI reviewers across five dimensions: Content Faithfulness, Argumentative Alignment, Focus Consistency, Question Constructiveness, and AI-Likelihood. Notably, we propose a Max-Recall strategy to accommodate valid expert disagreement and introduce a curated dataset of paper with high-confidence reviews, rigorously filtered to remove procedural noise. Extensive experiments demonstrate that while traditional n-gram metrics fail to reflect human preferences, our proposed text-centric metrics--particularly the recall of weakness arguments--correlate strongly with rating accuracy. These findings establish that aligning AI critique focus with human experts is a prerequisite for reliable automated scoring, offering a robust standard for future research.

📄 PDF Abstract BibTeX arXiv:2604.19502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generative AI for Synthetic Data Across Multiple Medical Modalities: A Systematic Review of Recent Developments and Challenges

2024-06-27 · Mahmoud Ibrahim, Yasmina Al Khalil, Sina Amirrajab, Chang Sun 외

This paper presents a comprehensive systematic review of generative models (GANs, VAEs, DMs, and LLMs) used to synthesize various medical data types, including imaging (dermoscopic, mammographic, ultrasound, CT, MRI, and…

BenchmarkingClinical Knowledge

Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviews

2025-09-16 · Chenye Zou, Xingyue Wen, Tianyi Hu, Qian Janice Wang 외 arxiv

Recent advances in large language models (LLMs) have opened the door to culture-aware language tasks. We introduce the novel problem of adapting wine reviews across Chinese and English, which goes beyond literal translat…

Review-based Recommender Systems: A Survey of Approaches, Challenges and Future Perspectives

2024-05-09 · Emrul Hasan, Mizanur Rahman, Chen Ding, Jimmy Xiangji Huang 외

Recommender systems play a pivotal role in helping users navigate an overwhelming selection of products and services. On online platforms, users have the opportunity to share feedback in various modes, including numerica…

NavigateRecommendation Systems

CPR: Leveraging LLMs for Topic and Phrase Suggestion to Facilitate Comprehensive Product Reviews

2025-04-18 · Ekta Gujral, Apurva Sinha, Lishi Ji, Bijayani Sanghamitra Mishra

Consumers often heavily rely on online product reviews, analyzing both quantitative ratings and textual descriptions to assess product quality. However, existing research hasn't adequately addressed how to systematically…

ReviewSense: Transforming Customer Review Dynamics into Actionable Business Insights

2025-10-18 · Siddhartha Krothapalli, Kartikey Singh Bhandari, Tridib Kumar Das, Praveen Kumar 외 arxiv

As customer feedback becomes increasingly central to strategic growth, the ability to derive actionable insights from unstructured reviews is essential. While traditional AI-driven systems excel at predicting user prefer…

Sentiment Analysis