paper-with-me

홈 › Papers

Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks

2025-09-28 · Aaron Xuxiang Tian, Ruofan Zhang, Jiayao Tang, Young Min Cho, Xueqian Li, Qiang Yi, Ji Wang, Zhunping Zhang, Danrui Qi, Zekun Li, Xingyu Xiang, Sharath Chandra Guntuku, Lyle Ungar, Tianyu Shi, Chi Wang arxiv

We study multi-turn multi-agent orchestration, where multiple large language model (LLM) agents interact over multiple turns by iteratively proposing answers or casting votes until reaching consensus. Using four LLMs (Gemini 2.5 Pro, GPT-5, Grok 4, and Claude Sonnet 4) on GPQA-Diamond, IFEval, and MuSR, we conduct two experiments: (i) benchmarking orchestration against single-LLM baselines; and (ii) ablations on GPQA-Diamond that vary whether agents see who authored answers and whether they can observe ongoing votes. Orchestration matches or exceeds the strongest single model and consistently outperforms the others. Analysis of best-achievable orchestration performance shows potential for further gains. The ablations show that revealing authorship increases self-voting and ties, and that showing ongoing votes amplifies herding, which speeds convergence but can sometimes yield premature consensus.

📄 PDF Abstract BibTeX arXiv:2509.23537

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Drives Interactive Improvement from Feedback?

2026-06-29 · Bartłomiej Cupiał, Jan Łojek, Mikołaj Garstecki, Szymon Pobłocki 외 arxiv

We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone. In multi-turn language agent setting, higher final accuracy can reflect useful feedback, but it can a…

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

2026-05-24 · Xiang Cheng, Yulan Hu, Lulu Zheng, Zheng Pan 외 arxiv

Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single user, where the field is approaching saturation. This single-user assumptio…

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

2026-06-01 · Rui Yang, Qianhui Wu, Yuxi Chen, Hao Bai 외 arxiv

Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the strongest systems remain largely proprietary, wh…

Reinforcement Learning

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

2026-06-05 · Yilun Liu, Miao Zhang, Shimin Tao, Minggui He 외 arxiv

Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and insight-poor, necessitating fine-grained multilingual post-evaluation …

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs

2026-05-09 · Wenzhi Fang, Liangqi Yuan, Guangchen Lan, Dong-Jun Han 외 arxiv

Multi-agent large language model (LLM) systems often rely on a controller to coordinate a pool of heterogeneous models, yet existing controllers are typically limited to one-shot routing: they select a model once and ret…