paper-with-me

홈 › Papers

Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

2026-02-10 · Wei Yang, Shixuan Li, Heng Ping, Peiyu Zhang, Paul Bogdan, Jesse Thomason arxiv

Multi-agent systems (MAS) can substantially extend the reasoning capacity of large language models (LLMs), yet most frameworks still aggregate agent outputs with majority voting. This heuristic discards the evidential structure of reasoning traces and is brittle under the confabulation consensus, where agents share correlated biases and converge on the same incorrect rationale. We introduce AgentAuditor, which replaces voting with a path search over a Reasoning Tree that explicitly represents agreements and divergences among agent traces. AgentAuditor resolves conflicts by comparing reasoning branches at critical divergence points, turning global adjudication into efficient, localized verification. We further propose Anti-Consensus Preference Optimization (ACPO), which trains the adjudicator on majority-failure cases and rewards evidence-based minority selections over popular errors. AgentAuditor is agnostic to MAS setting, and we find across 5 popular settings that it yields up to 5% absolute accuracy improvement over a majority vote, and up to 3% over using LLM-as-Judge.

📄 PDF Abstract BibTeX arXiv:2602.09341

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing

2026-05-08 · Yinsheng Yao, Jiehao Tang, Zhaozhen Yang, Dawei Cheng arxiv

While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verification, allowing early errors to cascade unchecked. This lack of modul…

TRUST: A Framework for Decentralized AI Service v.0.1

2026-04-29 · Yu-Chao Huang, Zhen Tan, Mohan Zhang, Pingzhi Li 외 arxiv

Large Reasoning Models (LRMs) and Multi-Agent Systems (MAS) in high-stakes domains demand reliable verification, yet centralized approaches suffer four limitations: (1) Robustness, with single points of failure vulnerabl…

AuditAgent: Expert-Guided Multi-Agent Reasoning for Cross-Document Fraudulent Evidence Discovery

2025-09-30 · Songran Bai, Bingzhe Wu, Yiwei Zhang, Chengke Wu 외 arxiv

Financial fraud detection in real-world scenarios presents significant challenges due to the subtlety and dispersion of evidence across complex, multi-year financial disclosures. In this work, we introduce a novel multi-…

Fraud Detection

Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data

2025-09-25 · George Yakushev, Alina Shutova, Ivan Rubachev, Natalia Bereberdina 외 arxiv

Tabular foundation models are becoming increasingly popular for low-resource tabular problems. These models make up for small training datasets by pretraining on large volumes of synthetic data. The prior knowledge obtai…

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

2026-04-09 · Wenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu 외 arxiv

In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However, coherent reasoning can still violate logical or evidential constrain…