paper-with-me

Papers

Benchmarking Agentic Review Systems

2026-06-18 · Dang Nguyen, Wanqing Hao, Yanai Elazar, Chenhao Tan arxiv

A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coarse), one proprietary system (Reviewer3), and a zero-shot baseline, across six LLMs spanning frontier and efficient models. First, we study whether AI reviews on ICLR/NeurIPS papers track with papers' quality as approximated by external signals such as citations and acceptance decisions. Every system performs above chance in pairwise accuracy, and the best is OpenAIReview + GPT-5.5 at 83.0%. Second, to test whether systems can catch errors with known ground truth, we construct a perturbation benchmark that injects four categories of errors into papers across eight arXiv subject classes and measure detection recall. The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors, leaving substantial room for improvement. The union of detections across six models reaches 83.3% recall, suggesting different models detect different errors and better harness design can potentially increase performance. Beyond these benchmarks, we study a public deployment of OpenAIReview with real users. Votes on its comments skew positive at 1.44 to 1, and the most common complaints are about false positives and minor nitpicks. Together, by evaluating full review systems backed by state-of-the-art models on real research papers, we show that while AI reviews still have room for improvement, they can already track human quality judgments well, catch important errors, and earn positive feedback from real users.

📄 PDF Abstract BibTeX arXiv:2606.19749

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agentic AI Systems in Electrical Power Systems Engineering: Current State-of-the-Art and Challenges

2025-11-18 · Soham Ghosh, Gaurav Mittal arxiv

Agentic AI systems have recently emerged as a critical and transformative approach in artificial intelligence, offering capabilities that extend far beyond traditional AI agents and contemporary generative AI models. Thi…

Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems

2025-03-09 · Dany Moshkovich, Hadar Mulian, Sergey Zeltyn, Natti Eder 외

The rise of agentic AI systems, where agents collaborate to perform diverse tasks, poses new challenges with observing, analyzing and optimizing their behavior. Traditional evaluation and benchmarking approaches struggle…

Benchmarking

Towards Agentic AI Governance: A Preliminary Assessment

2026-07-08 · Mubarak Raji, Masooda Bashir arxiv

Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely characterized as the Year of Agentic AI, 2025 marked accelerated development …

A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framework Development

2026-06-12 · Ahmed Mohammed Almalki, Mehedi Masud arxiv

This paper presents a structured analysis of security challenges in long-horizon agentic AI systems. The study reviews existing threats, evaluation approaches, attack propagation mechanisms, and security frameworks. A ta…

Detecting Silent Failures in Multi-Agentic AI Trajectories

2025-11-06 · Divya Pathak, Harshit Kumar, Anuska Roy, Felix George 외 arxiv

Multi-Agentic AI systems, powered by large language models (LLMs), are inherently non-deterministic and prone to silent failures such as drift, cycles, and missing details in outputs, which are difficult to detect. We in…

Anomaly Detection