paper-with-me

홈 › Papers

Multi-agent Policy Reciprocity with Theoretical Guarantee

2023-04-12 · Haozhi Wang, Yinchuan Li, Qing Wang, Yunfeng Shao, Jianye Hao

Modern multi-agent reinforcement learning (RL) algorithms hold great potential for solving a variety of real-world problems. However, they do not fully exploit cross-agent knowledge to reduce sample complexity and improve performance. Although transfer RL supports knowledge sharing, it is hyperparameter sensitive and complex. To solve this problem, we propose a novel multi-agent policy reciprocity (PR) framework, where each agent can fully exploit cross-agent policies even in mismatched states. We then define an adjacency space for mismatched states and design a plug-and-play module for value iteration, which enables agents to infer more precise returns. To improve the scalability of PR, deep PR is proposed for continuous control tasks. Moreover, theoretical analysis shows that agents can asymptotically reach consensus through individual perceived rewards and converge to an optimal value function, which implies the stability and effectiveness of PR, respectively. Experimental results on discrete and continuous environments demonstrate that PR outperforms various existing RL and transfer RL methods.

📄 PDF Abstract BibTeX arXiv:2304.05632

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous ControlMulti-agent Reinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Cooperation Through Indirect Reciprocity in Child-Robot Interactions

2025-11-07 · Isabel Neto, Alexandre S. Pires, Filipa Correia, Fernando P. Santos arxiv

Social interactions increasingly involve artificial agents, such as conversational or collaborative bots. Understanding trust and prosociality in these settings is fundamental to improve human-AI teamwork. Research in bi…

Best Response Shaping

2024-04-05 · Milad Aghajohari, Tim Cooijmans, Juan Agustin Duque, Shunichi Akatsuka 외

We investigate the challenge of multi-agent deep reinforcement learning in partially competitive environments, where traditional methods struggle to foster reciprocity-based cooperation. LOLA and POLA agents learn recipr…

Deep Reinforcement LearningQuestion Answering

Multi-Agent Guided Policy Optimization

2025-07-24 · Yueheng Li, Guangming Xie, Zongqing Lu arxiv

Due to practical constraints such as partial observability and limited communication, Centralized Training with Decentralized Execution (CTDE) has become the dominant paradigm in cooperative Multi-Agent Reinforcement Lea…

Multi-agent Reinforcement Learning

Symmetric (Optimistic) Natural Policy Gradient for Multi-agent Learning with Parameter Convergence

2022-10-23 · Sarath Pattathil, Kaiqing Zhang, Asuman Ozdaglar

Multi-agent interactions are increasingly important in the context of reinforcement learning, and the theoretical foundations of policy gradient methods have attracted surging research interest. We investigate the global…

Policy Gradient Methods

Off-Policy Correction For Multi-Agent Reinforcement Learning

2021-11-22 · Michał Zawalski, Błażej Osiński, Henryk Michalewski, Piotr Miłoś

Multi-agent reinforcement learning (MARL) provides a framework for problems involving multiple interacting agents. Despite apparent similarity to the single-agent case, multi-agent problems are often harder to train and …

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1