paper-with-me

홈 › Papers

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

2026-05-09 · Minzheng Wang, Run Luo, Yanbo Wang, Zichen Liu, Yuqiao Tan, Tao Tan, Xu Nan, Yinhe Zheng, Wenji Mao arxiv

While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a critical issue: evolution impasse. Due to the vast strategy space, language agents frequently converge to homogenized behaviors, leading to deterministic match outcomes that eliminate the gradient signals necessary for policy evolution. To tackle this issue, we propose Dual-scale Evolutionary Policy Training (DEPT) for social language games. DEPT introduces a time-scaled evolutionary perception mechanism that detects impasse by quantifying dual-scale value baseline divergence alongside match entropy. Upon perceiving the collapse, it then activates asymmetric advantage reshaping to dynamically modulate the optimization landscape for intervention. Thus, our method effectively restores gradient signals and enforces sustained strategic exploration. Extensive experiments on multiple social language games demonstrate that DEPT outperforms strong baselines, avoiding policy degeneration and driving the continuous evolution of social language agents.

📄 PDF Abstract BibTeX arXiv:2605.08721

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

MedGR$^2$: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning

2025-08-28 · Weihai Zhi, Jiayan Guo, Shangyang Li arxiv

The application of Vision-Language Models (VLMs) in medicine is critically hampered by the scarcity of high-quality, expert-annotated data. Supervised Fine-Tuning (SFT) on existing datasets often leads to poor generaliza…

Reinforcement Learning

ERL-Re$^2$: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy Representation

2022-10-26 · Jianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng 외

Deep Reinforcement Learning (Deep RL) and Evolutionary Algorithms (EA) are two major paradigms of policy optimization with distinct learning principles, i.e., gradient-based v.s. gradient-free. An appealing research dire…

continuous-controlContinuous ControlDeep Reinforcement LearningEvolutionary Algorithms+2

Breaking Task Impasses Quickly: Adaptive Neuro-Symbolic Learning for Open-World Robotics

2026-01-01 · Pierrick Lorang arxiv

Adapting to unforeseen novelties in open-world environments remains a major challenge for autonomous systems. While hybrid planning and reinforcement learning (RL) approaches show promise, they often suffer from sample i…

Reinforcement LearningAutonomous DrivingMotion Planning

An Ensemble of Evolutionary Algorithms With Both Crisscross Search and Sparrow Search for Processing Inferior Individuals

2026-01-15 · Mingxuan Du, Tingzhang Luo, Ziyang Wang, Chengjun Li arxiv

In the field of artificial intelligence, real parameter single objective optimization is an important direction. Both the Differential Evolution (DE) and the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) demon…

ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models

2025-11-17 · Siyang Cheng, Gaotian Liu, Rui Mei, Yilin Wang 외 arxiv

The rapid adoption of large language models (LLMs) has brought both transformative applications and new security risks, including jailbreak attacks that bypass alignment safeguards to elicit harmful outputs. Existing aut…