paper-with-me

홈 › Papers

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

2026-06-11 · Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee, Daniel Miao, Peter J. Gilbert, Nick Hynes, Mauro Staver, Warren He, David Marn, Andrew Low, Xi Zhang, Elron Bandel, Michal Shmueli-Scheuer, Siva Reddy, Alexandre Drouin, Alexandre Lacoste, Ramayya Krishnan, Elham Tabassi, Yu Su, Victor Barres, Chenguang Wang, Wenbo Guo, Dawn Song arxiv

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.

📄 PDF Abstract BibTeX arXiv:2606.13608

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Challenges in Credit Assignment for Multi-Agent Reinforcement Learning in Open Agent Systems

2025-10-31 · Alireza Saleh Abadi, Leen-Kiat Soh arxiv

In the rapidly evolving field of multi-agent reinforcement learning (MARL), understanding the dynamics of open systems is crucial. Openness in MARL refers to the dynam-ic nature of agent populations, tasks, and agent typ…

Multi-agent Reinforcement Learning

Agentifying Agentic AI

2025-11-21 · Virginia Dignum, Frank Dignum arxiv

Agentic AI seeks to endow systems with sustained autonomy, reasoning, and interaction capabilities. To realize this vision, its assumptions about agency must be complemented by explicit models of cognition, cooperation, …

PLATO: Pointer Learner for Agent and Task Openness

2026-07-27 · Alireza Saleh Abadi, Leen-Kiat Soh, Daniel Alan Redder, Adam Eck 외 arxiv

Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a …

Multi-agent Reinforcement LearningZero-shot GeneralizationGraph Neural Network

Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP

2026-02-11 · Zeynab Anbiaee, Mahdi Rabbani, Mansur Mirani, Gunjan Piya 외 arxiv

The rapid development of the AI agent communication protocols, including the Model Context Protocol (MCP), Agent2Agent (A2A), Agora, and Agent Network Protocol (ANP), is reshaping how AI agents communicate with tools, se…

Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent

2026-03-10 · Josef Hardi, Martin J. O'Connor, Marcos Martinez-Romero, Jean G. Rosario 외 arxiv

Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reuse. Even when standard metadata reporting guidelines exist, they typically lack m…