paper-with-me

Papers

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems

2025-09-11 · Minghang Zhu, Zhengliang Shi, Zhiwei Xu, Shiguang Wu, Lingjie Wang, Pengjie Ren, Zhaochun Ren, Zhumin Chen arxiv

The advancement of large language models (LLMs) has enabled the construction of multi-agent systems to solve complex tasks by dividing responsibilities among specialized agents, such as a planning agent for subgoal generation and a grounding agent for executing tool-use actions. Most existing methods typically fine-tune these agents independently, leading to capability gaps among them with poor coordination. To address this, we propose MOAT, a Multi-Agent Joint Alignment Tuning framework that improves agents collaboration through iterative alignment. MOAT alternates between two key stages: (1) Planning Agent Alignment, which optimizes the planning agent to generate subgoal sequences that better guide the grounding agent; and (2) Grounding Agent Improving, which fine-tunes the grounding agent using diverse subgoal-action pairs generated by the agent itself to enhance its generalization capablity. Theoretical analysis proves that MOAT ensures a non-decreasing and progressively convergent training process. Experiments across six benchmarks demonstrate that MOAT outperforms state-of-the-art baselines, achieving average improvements of 3.1% on held-in tasks and 4.4% on held-out tasks.

📄 PDF Abstract BibTeX arXiv:2509.09629

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

2025-08-06 · Tongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui Liu arxiv

Monocular and stereo depth estimation offer complementary strengths: monocular methods capture rich contextual priors but lack geometric precision, while stereo approaches leverage epipolar geometry yet struggle with amb…

Zero-shot GeneralizationStereo Depth Estimation

HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models

2025-11-27 · Haoxi Zeng, Haoxuan Li, Yi Bin, Pengpeng Zeng 외 arxiv

Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due to the lack of region-level supervision,…

Harmonizing word alignments and syntactic structures for extracting phrasal translation equivalents

2015-06-01 · WS 2015 6 · Dun Deng, Nianwen Xue, Shiman Guo
Machine TranslationTranslationWord Alignment

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

2025-07-29 · Shaojun E, Yuchen Yang, Jiaheng Wu, Yan Zhang 외 arxiv

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal mo…

MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation

2026-05-07 · Hao Wang, Shiqi Wang, Qi Liu arxiv

Generating realistic 3D Human-Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high-level semantic intent with strict low-l…