paper-with-me

Papers

M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality

2025-03-03 · Ziyan Wang, Zhicheng Zhang, Fei Fang, Yali Du

Designing effective reward functions in multi-agent reinforcement learning (MARL) is a significant challenge, often leading to suboptimal or misaligned behaviors in complex, coordinated environments. We introduce Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality ($\text{M}^3\text{HF}$), a novel framework that integrates multi-phase human feedback of mixed quality into the MARL training process. By involving humans with diverse expertise levels to provide iterative guidance, $\text{M}^3\text{HF}$ leverages both expert and non-expert feedback to continuously refine agents' policies. During training, we strategically pause agent learning for human evaluation, parse feedback using large language models to assign it appropriately and update reward functions through predefined templates and adaptive weights by using weight decay and performance-based adjustments. Our approach enables the integration of nuanced human insights across various levels of quality, enhancing the interpretability and robustness of multi-agent cooperation. Empirical results in challenging environments demonstrate that $\text{M}^3\text{HF}$ significantly outperforms state-of-the-art methods, effectively addressing the complexities of reward design in MARL and enabling broader human participation in the training process.

📄 PDF Abstract BibTeX arXiv:2503.02077

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Meta-Reinforcement Learning Using Model Parameters

2022-10-27 · Gabriel Hartmann, Amos Azaria

In meta-reinforcement learning, an agent is trained in multiple different environments and attempts to learn a meta-policy that can efficiently adapt to a new environment. This paper presents RAMP, a Reinforcement learni…

Meta Reinforcement Learningmodelreinforcement-learningReinforcement Learning+1

Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning

2019-02-18 · Woojun Kim, Myungsik Cho, Youngchul Sung

In this paper, we propose a new learning technique named message-dropout to improve the performance for multi-agent deep reinforcement learning under two application scenarios: 1) classical multi-agent reinforcement lear…

Deep Reinforcement LearningMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+1

Human-Centric Traffic Signal Control for Equity: A Multi-Agent Action Branching Deep Reinforcement Learning Approach

2026-02-03 · Xiaocai Zhang, Neema Nassir, Lok Sang Chan, Milad Haghani arxiv

Coordinating traffic signals along multimodal corridors is challenging because many multi-agent deep reinforcement learning (DRL) approaches remain vehicle-centric and struggle with high-dimensional discrete action space…

Reinforcement Learning

Emergent Tool Use From Multi-Agent Autocurricula

2019-09-17 · ICLR 2020 1 · Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu 외

Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct roun…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Emergent Coordination and Phase Structure in Independent Multi-Agent Reinforcement Learning

2025-11-28 · Azusa Yamaguchi arxiv

A clearer understanding of when coordination emerges, fluctuates, or collapses in decentralized multi-agent reinforcement learning (MARL) is increasingly sought in order to characterize the dynamics of multi-agent learni…

Multi-agent Reinforcement Learning