paper-with-me

Papers

Towards Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games

2021-12-01 · NeurIPS 2021 12 · Xiangyu Liu, Hangtian Jia, Ying Wen, Yaodong Yang, Yujing Hu, Yingfeng Chen, Changjie Fan, Zhipeng Hu

Measuring and promoting policy diversity is critical for solving games with strong non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). With that in mind, maintaining a pool of diverse policies via open-ended learning is an attractive solution, which can generate auto-curricula to avoid being exploited. However, in conventional open-ended learning algorithms, there are no widely accepted definitions for diversity, making it hard to construct and evaluate the diverse policies. In this work, we summarize previous concepts of diversity and work towards offering a unified measure of diversity in multi-agent open-ended learning to include all elements in Markov games, based on both Behavioral Diversity (BD) and Response Diversity (RD). At the trajectory distribution level, we re-define BD in the state-action space as the discrepancies of occupancy measures. For the reward dynamics, we propose RD to characterize diversity through the responses of policies when encountering different opponents. We also show that many current diversity measures fall in one of the categories of BD or RD but not both. With this unified diversity measure, we design the corresponding diversity-promoting objective and population effectivity when seeking the best responses in open-ended learning. We validate our methods in both relatively simple games like matrix game, non-transitive mixture model, and the complex \textit{Google Research Football} environment. The population found by our methods reveals the lowest exploitability, highest population effectivity in matrix game and non-transitive mixture model, as well as the largest goal difference when interacting with opponents of various levels in \textit{Google Research Football}.

📄 PDF Abstract BibTeX

Code (1)

sjtu-marl/bd_rd_psro 공식 구현 pytorch

Tasks

Diversity

Similar Papers 제목 키워드 기반

Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games

2021-06-09 · Xiangyu Liu, Hangtian Jia, Ying Wen, Yaodong Yang 외

Measuring and promoting policy diversity is critical for solving games with strong non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). With that in mind, m…

Diversity

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation

2026-05-18 · Guining Cao, Jiaxin Peng, Chu Zeng, Yu Zhao 외 arxiv

Current reinforcement learning(RL) methods are broadly applicable and powerful in verifiable settings where scalar rewards can be provided. However, in open-ended generation tasks, verifying the correctness of responses …

Reinforcement Learning

AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity Optimization

2025-06-05 · Saeed Hedayatian, Stefanos Nikolaidis

Quality-Diversity (QD) algorithms have shown remarkable success in discovering diverse, high-performing solutions, but rely heavily on hand-crafted behavioral descriptors that constrain exploration to predefined notions …

continuous-controlContinuous ControlDiversitySequential Decision Making+1

Modelling Behavioural Diversity for Learning in Open-Ended Games

2021-03-14 · Nicolas Perez Nieves, Yaodong Yang, Oliver Slumbers, David Henry Mguni 외

Promoting behavioural diversity is critical for solving games with non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). Yet, there is a lack of rigorous tre…

DiversityPoint Processes

Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization

2023-10-18 · Li Ding, Jenny Zhang, Jeff Clune, Lee Spector 외

Reinforcement Learning from Human Feedback (RLHF) has shown potential in qualitative tasks where easily defined performance measures are lacking. However, there are drawbacks when RLHF is commonly used to optimize for av…

DiversityImage Generationreinforcement-learningReinforcement Learning+4