paper-with-me

홈 › Papers

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

2025-10-13 · Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir, Zachary S Siegel, Boyi Wei, Tianci Xue, Ziru Chen, Felix Chen, Saiteja Utpala, Franck Ndzomga, Dheeraj Oruganty, Sophie Luskin, Kangheng Liu, Botao Yu, Amit Arora, Dongyoon Hahm, Harsh Trivedi, Huan Sun, Juyong Lee, Tengjun Jin, Yifan Mai, Yifei Zhou, Yuxuan Zhu, Rishi Bommasani, Daniel Kang, Dawn Song, Peter Henderson, Yu Su, Percy Liang, Arvind Narayanan arxiv

AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of how well agents really work. We introduce the Holistic Agent Leaderboard (HAL) to address these challenges. We make three main contributions. First, we provide a standardized evaluation harness that orchestrates parallel evaluations across hundreds of VMs, reducing evaluation time from weeks to hours while eliminating common implementation bugs. Second, we conduct three-dimensional analysis spanning models, scaffolds, and benchmarks. We validate the harness by conducting 21,730 agent rollouts across 9 models and 9 benchmarks in coding, web navigation, science, and customer service with a total cost of about $40,000. Our analysis reveals surprising insights, such as higher reasoning effort reducing accuracy in the majority of runs. Third, we use LLM-aided log inspection to uncover previously unreported behaviors, such as searching for the benchmark on HuggingFace instead of solving a task, or misusing credit cards in flight booking tasks. We share all agent logs, comprising 2.5B tokens of language model calls, to incentivize further research into agent behavior. By standardizing how the field evaluates agents and addressing common pitfalls in agent evaluation, we hope to shift the focus from agents that ace benchmarks to agents that work reliably in the real world.

📄 PDF Abstract BibTeX arXiv:2510.11977

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

2026-07-31 · Yuan Gao, Zeren Yang, Junnan Li, Shawn 외 arxiv

Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, b…

Agentic Time Machine as an Infrastructure for Future-Event Forecasting

2026-06-19 · Jingyi Chai, Bingyang Zheng, Xiangrui Liu, Hao Lu 외 arxiv

Forecasting future events is a critical challenge for large language model (LLM) agents, spanning domains from elections and monetary policy to financial markets. However, evaluating progress on this task presents a fund…

APEX-Agents

2026-01-20 · Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly 외 arxiv

We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultant…

Attacking Autonomous Driving Agents with Adversarial Machine Learning: A Holistic Evaluation with the CARLA Leaderboard

2025-11-18 · Henry Wong, Clement Fung, Weiran Lin, Karen Li 외 arxiv

To autonomously control vehicles, driving agents use outputs from a combination of machine-learning (ML) models, controller logic, and custom modules. Although numerous prior works have shown that adversarial examples ca…

Autonomous Driving

Integration of SCADA services in cross-infrastructure holistic tests of cyber-physical energy systems

2019-05-15

Cyber-Physical Energy System, due to its multi-domain nature, requires a holistic validation methodology, which may involve the integration of assets and expertise from various research infrastructures. In this paper, th…