paper-with-me

홈 › Papers

Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

2026-02-09 · Yanming Li, Xuelin Zhang, WenJie Lu, Ziye Tang, Maodong Wu, Haotian Luo, Tongtong Wu, Zijie Peng, Hongze Mi, Yibo Feng, Naiqiang Tan, Chao Huang, Lian Peng, Li Shen arxiv

Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems. However, training these systems remains notoriously difficult due to the credit assignment challenge, as it is often unclear which specific functional agent is responsible for the success or failure of decision trajectories. Existing methods typically rely on sparse or globally broadcast rewards, failing to capture individual contributions and leading to inefficient reinforcement learning. To address these limitations, we introduce the Shapley-based Hierarchical Attribution for Reinforcement Policy (SHARP), a novel framework for optimizing multi-agent reinforcement learning via precise credit attribution. SHARP effectively stabilizes training by normalizing agent-specific advantages across trajectory groups, primarily through a decomposed reward mechanism comprising a global broadcast-accuracy reward, a Shapley-based marginal-credit reward for each agent, and a tool-process reward to improve execution efficiency. Extensive experiments across various real-world benchmarks demonstrate that SHARP significantly outperforms recent state-of-the-art baselines, achieving average match improvements of 23.66% and 14.05% over single-agent and multi-agent approaches, respectively.

📄 PDF Abstract BibTeX arXiv:2602.08335

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

Shapley Q-value: A Local Reward Approach to Solve Global Reward Games

2019-07-11 · Jianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie Gu

Cooperative game is a critical research area in the multi-agent reinforcement learning (MARL). Global reward game is a subclass of cooperative games, where all agents aim to maximize the global reward. Credit assignment …

Multi-agent Reinforcement LearningPolicy Gradient MethodsReinforcement Learning

Owen-Shapley Policy Optimization: A Principled RL Algorithm for Generative Search LLMs

2026-01-13 · Abhijnan Nath, Alireza Bagheri Garakani, Tianchen Zhou, Fan Yang 외 arxiv

Large language models are increasingly trained via reinforcement learning for personalized recommendation tasks, but standard methods like GRPO rely on sparse, sequence-level rewards. These obscure which tokens actually …

Reinforcement Learning

Shapley Value Based Multi-Agent Reinforcement Learning: Theory, Method and Its Application to Energy Network

2024-02-23 · Jianhong Wang

Multi-agent reinforcement learning is an area of rapid advancement in artificial intelligence and machine learning. One of the important questions to be answered is how to conduct credit assignment in a multi-agent syste…

Learning TheoryMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning

STAS: Spatial-Temporal Return Decomposition for Multi-agent Reinforcement Learning

2023-04-15 · Sirui Chen, Zhaowei Zhang, Yaodong Yang, Yali Du

Centralized Training with Decentralized Execution (CTDE) has been proven to be an effective paradigm in cooperative multi-agent reinforcement learning (MARL). One of the major challenges is credit assignment, which aims …

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning

Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning

2021-06-01 · Jiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu 외

Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challe…

counterfactualMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+3