paper-with-me

홈 › Papers

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees

2026-04-17 · Yi Xie, Yangyang Xu, Yi Fan, Bo Liu arxiv

Large language models (LLMs) with a large number of parameters achieve strong performance but are often prohibitively expensive to deploy. Recent work explores using teams of smaller, more efficient LLMs that collectively match or even outperform a single large model. However, jointly updating multiple agents introduces compounding distribution shifts, making coordination and stability during training difficult. We address this by introducing Sequential Agent Tuning (SAT), a coordinator-free training paradigm. SAT represents the team as a factorized policy and employs block-coordinate updates over agents, enabling scalable, decentralized training without a central controller. Specifically, we develop a sequence-aware, on-policy advantage estimator that conditions on the evolving team policy, coupled with per-agent KL trust regions that isolate occupancy drift. Theoretically, this framework provides two critical guarantees. First, it ensures monotonic improvement, stabilizing the training process. Second, it establishes provable plug-and-play invariance: any agent can be upgraded to a stronger model without retraining the rest of the team, with a formal guarantee that the performance bound improves. Empirically, a team of three 4B agents (12B total) trained with SAT surpasses the much larger Qwen3-32B on AIME24/25 benchmarks by 3.9\% on average. We validate our plug-and-play theory by swapping in two 8B agents, which boosts the composite score by 10.4\%. We provide code and appendix of proof at https://github.com/Yydc/SAT-AAMAS

📄 PDF Abstract BibTeX arXiv:2605.05216

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Design of a GIS-based Assistant Software Agent for the Incident Commander to Coordinate Emergency Response Operations

2014-01-01 · Reza Nourjou, Michinori Hatayama, Stephen F. Smith, Atabak Sadeghi 외

Problem: This paper addresses the design of an intelligent software system for the IC (incident commander) of a team in order to coordinate actions of agents (field units or robots) in the domain of emergency/crisis resp…

Beyond Relevance: Bayesian Evidence Acquisition for Agentic Whole-Slide Image Reasoning

2026-08-06 · Bryan Wong, Xun Xu, Huazhu Fu, Nancy F. Chen 외 arxiv

Whole-slide image (WSI) reasoning requires an agent to sequentially acquire visual evidence before answering a diagnostic question. Existing training-free agentic frameworks formulate this process as iterative patch retr…

Deep-Unfolded Coordination

2026-06-18 · Hunter Kuperman, Minchan Jung, Rahul V. Ghosh, Alex Oshin 외 arxiv

Distributed optimization is a highly scalable and structurally transparent technique to solve multi-agent robotics problems; however, such methods often suffer from the need for highly-specialized, problem-specific hyper…

Distributed Optimization

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination

2026-05-01 · Yi Xie, Siao Liu, Falong Fan, Yuanqi Yao 외 arxiv

Multi-agent LLM systems have shown promise for complex reasoning, yet recent evaluations reveal they often underperform single-model baselines. We identify a structural failure mode in sequential fine-tuning of shared-co…

Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation

2025-11-27 · Zehao Deng, Tianjie Ju, Zheng Wu, Zhuosheng Zhang 외 arxiv

The rapid development of large vision-language model (VLM) has greatly promoted the research of GUI agent. However, GUI agents still face significant challenges in handling long-horizon tasks. First, single-agent models …

Reinforcement Learning