paper-with-me

Papers

Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation

2026-03-26 · Yeonjun In, Mehrab Tanjim, Jayakumar Subramanian, Sungchul Kim, Uttaran Bhattacharya, Wonjoong Kim, Sangwu Park, Somdeb Sarkhel, Chanyoung Park arxiv

Failure attribution is essential for diagnosing and improving multi-agent systems (MAS), yet existing benchmarks and methods largely assume a single deterministic root cause for each failure. In practice, MAS failures often admit multiple plausible attributions due to complex inter-agent dependencies and ambiguous execution trajectories. We revisit MAS failure attribution from a multi-perspective standpoint and propose multi-perspective failure attribution, a practical paradigm that explicitly accounts for attribution ambiguity. To support this setting, we introduce MP-Bench, the first benchmark designed for multi-perspective failure attribution in MAS, along with a new evaluation protocol tailored to this paradigm. Through extensive experiments, we find that prior conclusions suggesting LLMs struggle with failure attribution are largely driven by limitations in existing benchmark designs. Our results highlight the necessity of multi-perspective benchmarks and evaluation protocols for realistic and reliable MAS debugging.

📄 PDF Abstract BibTeX arXiv:2603.25001

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

2025-04-30 · Shaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu 외

Failure attribution in LLM multi-agent systems-identifying the agent and step responsible for task failures-provides crucial clues for systems debugging but remains underexplored and labor-intensive. In this paper, we pr…

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

2026-05-17 · Hezhe Qiao, Hanghang Tong, Ee-Peng Lim, Bing Liu 외 arxiv

Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliability. Automatic failure attribution is therefore critical, but existi…

DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

2026-09-04 · Zehao Wang, Lanjun Wang, Shilong Jin, Junjie Chen 외 arxiv

Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lea…

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems

2026-06-02 · Taiyu Zhu, Yifan Wu, Weilin Jin, Ying Li 외 arxiv

LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks. However, these systems are highly sensitive to single-step execution errors that can propagate through agent intera…

Text Generation

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference

2025-09-10 · Guoqing Ma, Jia Zhu, Hanghui Guo, Weijie Shi 외 arxiv

Multi-agent systems (MAS) are critical for automating complex tasks, yet their practical deployment is severely hampered by the challenge of failure attribution. Current diagnostic tools, which rely on statistical correl…

Causal Inference