paper-with-me

홈 › Papers

Expert-Space Exploration in MoE Reinforcement Learning

2026-09-11 · Hongyi He, Zhenghao Lin, Xiao Liu, Peng Cheng, Yan Lu, Yeyun Gong hf

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the sparse computation paths that induce output distributions, expert selection offers an additional source of rollout diversity. Through empirical analysis, we find that perturbing expert routing effectively alters model output and increases rollout diversity, which is similar to increasing the decoding temperature. However, direct perturbation can activate unsuitable experts and substantially degrade rollout quality. Motivated by these observations, we introduce Expert-Space Exploration Reinforcement Learning (ESRL), an architecture-aware framework that explicitly explores the expert-routing space of MoE models. ESRL preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths. The perturbation strength is further adapted according to router entropy to avoid over-perturbation. To mitigate the routing mismatch introduced by perturbation, ESRL records the expert paths used during rollout and replays them during policy optimization. Experiments demonstrate that ESRL achieves the best performance across MoE backbones with top-K, top-1, and shared-expert routing, as well as across mathematics, science, and code tasks without additional sampling or computational cost. Specifically, ESRL on Qwen3-30B-A3B achieves the best among all compared methods, improving average Pass@1 and Pass@8 over GRPO by 3.2 and 4.5 percentage points, respectively. Further analyses of expert utilization and training dynamics provide insights into how exploiting MoE-specific routing structure benefits RL training.

📄 PDF Abstract BibTeX arXiv:2609.13058

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Robust Exploration in Directed Controller Synthesis via Reinforcement Learning with Soft Mixture-of-Experts

2026-02-22 · Toshihide Ubukata, Zhiyao Wang, Enhong Mu, Jialong Li 외 arxiv

On-the-fly Directed Controller Synthesis (OTF-DCS) mitigates state-space explosion by incrementally exploring the system and relies critically on an exploration policy to guide search efficiently. Recent reinforcement le…

Zero-shot GeneralizationReinforcement Learning

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration

2025-06-25 · Heyang Zhao, Xingrui Yu, David M. Bossens, Ivor W. Tsang 외

Imitation learning is a central problem in reinforcement learning where the goal is to learn a policy that mimics the expert's behavior. In practice, it is often challenging to learn the expert policy from a limited numb…

Imitation LearningMuJoCo

Learning to Drive Using Sparse Imitation Reinforcement Learning

2022-05-24 · Yuci Han, Alper Yilmaz

In this paper, we propose Sparse Imitation Reinforcement Learning (SIRL), a hybrid end-to-end control policy that combines the sparse expert driving knowledge with reinforcement learning (RL) policy for autonomous drivin…

Autonomous Drivingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning

2025-12-15 · Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui 외 arxiv

Current Vision-Language-Action (VLA) paradigms in autonomous driving primarily rely on Imitation Learning (IL), which introduces inherent challenges such as distribution shift and causal confusion. Online Reinforcement L…

Reinforcement LearningAutonomous Driving

Mixture of Autoencoder Experts Guidance using Unlabeled and Incomplete Data for Exploration in Reinforcement Learning

2025-07-21 · Elias Malomgré, Pieter Simoens arxiv

Recent trends in Reinforcement Learning (RL) highlight the need for agents to learn from reward-free interactions and alternative supervision signals, such as unlabeled or incomplete demonstrations, rather than relying s…

Reinforcement Learning