paper-with-me

Papers

Survey on Evaluation of LLM-based Agents

2025-03-20 · Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, Michal Shmueli-Scheuer

The emergence of LLM-based agents represents a paradigm shift in AI, enabling autonomous systems to plan, reason, use tools, and maintain memory while interacting with dynamic environments. This paper provides the first comprehensive survey of evaluation methodologies for these increasingly capable agents. We systematically analyze evaluation benchmarks and frameworks across four critical dimensions: (1) fundamental agent capabilities, including planning, tool use, self-reflection, and memory; (2) application-specific benchmarks for web, software engineering, scientific, and conversational agents; (3) benchmarks for generalist agents; and (4) frameworks for evaluating agents. Our analysis reveals emerging trends, including a shift toward more realistic, challenging evaluations with continuously updated benchmarks. We also identify critical gaps that future research must address-particularly in assessing cost-efficiency, safety, and robustness, and in developing fine-grained, and scalable evaluation methods. This survey maps the rapidly evolving landscape of agent evaluation, reveals the emerging trends in the field, identifies current limitations, and proposes directions for future research.

📄 PDF Abstract BibTeX arXiv:2503.16416

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

2024-12-05 · Xiachong Feng, Longxu Dou, Ella Li, Qinghao Wang 외

Game-theoretic scenarios have become pivotal in evaluating the social intelligence of Large Language Model (LLM)-based social agents. While numerous studies have explored these agents in such settings, there is a lack of…

Language ModelingLanguage ModellingLarge Language ModelSurvey

SurveyBench: Can LLM(-Agents) Write Academic Surveys that Align with Reader Needs?

2025-10-03 · Zhaojun Sun, Xuzhou Zhu, Xuanhe Zhou, Xin Tong 외 arxiv

Academic survey writing, which distills vast literature into a coherent and insightful narrative, remains a labor-intensive and intellectually demanding task. While recent approaches, such as general DeepResearch agents …

SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation

2026-02-11 · Beichen Guo, Zhiyuan Wen, Jia Gu, Haochen Shi 외 arxiv

Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialized ASG frameworks and Deep Research agent…

Towards Trustworthy GUI Agents: A Survey

2025-03-30 · Yucheng Shi, Wenhao Yu, Wenlin Yao, Wenhu Chen 외

GUI agents, powered by large foundation models, can interact with digital interfaces, enabling various applications in web automation, mobile navigation, and software testing. However, their increasing autonomy has raise…

Decision MakingSequential Decision Makingsoftware testingSurvey

Evaluation and Benchmarking of LLM Agents: A Survey

2025-07-29 · Mahmoud Mohammadi, Yipeng Li, Jane Lo, Wendy Yip arxiv

The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agen…