paper-with-me

Papers

Role Play: Learning Adaptive Role-Specific Strategies in Multi-Agent Interactions

2024-11-02 · Weifan Long, Wen Wen, Peng Zhai, Lihua Zhang

Zero-shot coordination problem in multi-agent reinforcement learning (MARL), which requires agents to adapt to unseen agents, has attracted increasing attention. Traditional approaches often rely on the Self-Play (SP) framework to generate a diverse set of policies in a policy pool, which serves to improve the generalization capability of the final agent. However, these frameworks may struggle to capture the full spectrum of potential strategies, especially in real-world scenarios that demand agents balance cooperation with competition. In such settings, agents need strategies that can adapt to varying and often conflicting goals. Drawing inspiration from Social Value Orientation (SVO)-where individuals maintain stable value orientations during interactions with others-we propose a novel framework called \emph{Role Play} (RP). RP employs role embeddings to transform the challenge of policy diversity into a more manageable diversity of roles. It trains a common policy with role embedding observations and employs a role predictor to estimate the joint role embeddings of other agents, helping the learning agent adapt to its assigned role. We theoretically prove that an approximate optimal policy can be achieved by optimizing the expected cumulative reward relative to an approximate role-based policy. Experimental results in both cooperative (Overcooked) and mixed-motive games (Harvest, CleanUp) reveal that RP consistently outperforms strong baselines when interacting with unseen agents, highlighting its robustness and adaptability in complex environments.

📄 PDF Abstract BibTeX arXiv:2411.01166

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMulti-agent Reinforcement LearningRole Embedding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

RPO-PDT: Demonstrating Role-Play-Based Knowledge Adaptation for Student Support Dialogue (Demonstration System)

2026-06-08 · Filip Janik, Ewa Olton, Robert Smales, Harris Spratt 외 arxiv

We present RPO-PDT: a retrieval-grounded, role-play-based dialogue system for adaptive student support in higher education. RPO-PDT is: (1) able to provide institution-specific Personal Development Tutor (PDT) guidance u…

HyCoRA: Hyper-Contrastive Role-Adaptive Learning for Role-Playing

2025-11-11 · Shihao Yang, Zhicong Lu, Yong Yang, Bo Lv 외 arxiv

Multi-character role-playing aims to equip models with the capability to simulate diverse roles. Existing methods either use one shared parameterized module across all roles or assign a separate parameterized module to e…

Contrastive Learning

ChARM: Character-based Act-adaptive Reward Modeling for Advanced Role-Playing Language Agents

2025-05-29 · Feiteng Fang, Ting-En Lin, Yuchuan Wu, Xiong Liu 외

Role-Playing Language Agents (RPLAs) aim to simulate characters for realistic and engaging human-computer interactions. However, traditional reward models often struggle with scalability and adapting to subjective conver…

Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents

2025-09-17 · Xueqiao Zhang, Chao Zhang, Jingtao Xu, Yifan Zhu 외 arxiv

Role-playing agents (RPAs) have attracted growing interest for their ability to simulate immersive and interactive characters. However, existing approaches primarily focus on static role profiles, overlooking the dynamic…

Reasoning Does Not Necessarily Improve Role-Playing Ability

2025-02-24 · Xiachong Feng, Longxu Dou, Lingpeng Kong

The application of role-playing large language models (LLMs) is rapidly expanding in both academic and commercial domains, driving an increasing demand for high-precision role-playing models. Simultaneously, the rapid ad…