paper-with-me

홈 › Papers

DR-Arena: an Automated Evaluation Framework for Deep Research Agents

2026-01-15 · Yiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan Zhang arxiv

As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task performance has become a critical bottleneck. Current benchmarks predominantly rely on static datasets, which suffer from several limitations: limited task generality, temporal misalignment, and data contamination. To address these, we introduce DR-Arena, a fully automated evaluation framework that pushes DR agents to their capability limits through dynamic investigation. DR-Arena constructs real-time Information Trees from fresh web trends to ensure the evaluation rubric is synchronized with the live world state, and employs an automated Examiner to generate structured tasks testing two orthogonal capabilities: Deep reasoning and Wide coverage. DR-Arena further adopts Adaptive Evolvement Loop, a state-machine controller that dynamically escalates task complexity based on real-time performance, demanding deeper deduction or wider aggregation until a decisive capability boundary emerges. Experiments with six advanced DR agents demonstrate that DR-Arena achieves a Spearman correlation of 0.94 with the LMSYS Search Arena leaderboard. This represents the state-of-the-art alignment with human preferences without any manual efforts, validating DR-Arena as a reliable alternative for costly human adjudication.

📄 PDF Abstract BibTeX arXiv:2601.10504

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A3: Android Agent Arena for Mobile GUI Agents

2025-01-02 · Yuxiang Chai, Hanhao Li, Jiayu Zhang, Liang Liu 외

AI agents have become increasingly prevalent in recent years, driven by significant advancements in the field of large language models (LLMs). Mobile GUI agents, a subset of AI agents, are designed to autonomously perfor…

Information Retrieval

Understanding the Weakness of Large Language Model Agents within a Complex Android Environment

2024-02-09 · Mingzhe Xing, Rongkai Zhang, Hui Xue, Qi Chen 외

Large language models (LLMs) have empowered intelligent agents to execute intricate tasks within domain-specific software such as browsers and games. However, when applied to general-purpose software systems like operati…

Date UnderstandingLanguage ModelingLanguage ModellingLarge Language Model

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

2026-07-21 · Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz 외 arxiv

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agen…

SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

2025-07-01 · Yilun Zhao, Kaiyan Zhang, Tiansheng Hu, Sihong Wu 외 arxiv

We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciAr…

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

2026-06-10 · Tianyu Liu, Allen Xin Wang, Antonia Panescu, Lisa Xinyi Chen 외 arxiv

AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents rarely capture the com…