paper-with-me

홈 › Papers

What Is Your Agent's GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment

2025-10-09 · Allison Sihan Jia, Daniel Huang, Nikhil Vytla, Seung Won Wilson Yoo, Nirvika Choudhury, Shayak Sen, John C. Mitchell, Anupam Datta arxiv

We introduce the Agent GPA (Goal-Plan-Action) framework, driven by the fundamental insight that critical agent failures emerge at the intersections of setting goals, devising plans, and executing actions. We operationalize the framework with a factorized suite of LLM judges designed to measure distinct elements of Goal-Plan-Act alignment. To make this methodology scalable and generalizable across diverse agent architectures and datasets, we use state-of-the-art automated prompt optimization techniques to systematically generate domain-specific evaluation criteria. We validate this approach across three benchmarks: a multi-agent research setting (TRAIL/GAIA), a single coding agent setting (TRAIL/SWE-bench), and a private, enterprise data-agent setting (Snowflake Intelligence). Extensive evaluation on TRAIL/GAIA demonstrates the core validity of the framework, which identifies a broad range of agent failures (95% of human-annotated errors), localizes errors to enable targeted debugging (86% of human-annotated errors), and exhibits strong agreement with human evaluators. Crucially, by applying our automated methodology to both public datasets, we demonstrate that our GPA judges generally achieve the highest error coverage (ranging from 76% to 86%) in comparison to manual prompting approaches. We also leverage an evolutionary coding agent to improve judge consistency by up to 38% through iterative refinement of evaluation rubrics. Overall, Agent GPA provides a rigorous and generalizable paradigm for targeted agent evaluation.

📄 PDF Abstract BibTeX arXiv:2510.08847

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating and Modelling Hanabi-Playing Agents

2017-04-24 · Joseph Walton-Rivers, Piers R. Williams, Richard Bartle, Diego Perez-Liebana 외

Agent modelling involves considering how other agents will behave, in order to influence your own actions. In this paper, we explore the use of agent modelling in the hidden-information, collaborative card game Hanabi. W…

Game of Hanabi

Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents

2026-07-22 · Or Zion Eliav, Eyal Lenga, Shir Bernstien, Yisroel Mirsky arxiv

Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same treatment. We formalize agent reconnaissan…

EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?

2026-09-03 · Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li 외 hf

Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are…

Do Phone-Use Agents Respect Your Privacy?

2026-04-01 · Zhengyang Tang, Ke Ji, Xidong Wang, Zihan Ye 외 arxiv

We study whether phone-use agents respect privacy while completing benign mobile tasks. This question has remained hard to answer because privacy-compliant behavior is not operationalized for phone-use agents, and ordina…

Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents

2026-02-12 · Angelica Lim, Ö. Nilay Yalçin arxiv

The fields of human-robot interaction (HRI) and embodied conversational agents (ECAs) have long studied how empathy could be implemented in machines. One of the major drivers has been the goal of giving multimodal social…

Emotional Intelligence