paper-with-me

홈 › Papers

Robust Exploration in Directed Controller Synthesis via Reinforcement Learning with Soft Mixture-of-Experts

2026-02-22 · Toshihide Ubukata, Zhiyao Wang, Enhong Mu, Jialong Li, Kenji Tei arxiv

On-the-fly Directed Controller Synthesis (OTF-DCS) mitigates state-space explosion by incrementally exploring the system and relies critically on an exploration policy to guide search efficiently. Recent reinforcement learning (RL) approaches learn such policies and achieve promising zero-shot generalization from small training instances to larger unseen ones. However, a fundamental limitation is anisotropic generalization, where an RL policy exhibits strong performance only in a specific region of the domain-parameter space while remaining fragile elsewhere due to training stochasticity and trajectory-dependent bias. To address this, we propose a Soft Mixture-of-Experts framework that combines multiple RL experts via a prior-confidence gating mechanism and treats these anisotropic behaviors as complementary specializations. The evaluation on the Air Traffic benchmark shows that Soft-MoE substantially expands the solvable parameter space and improves robustness compared to any single expert.

📄 PDF Abstract BibTeX arXiv:2602.19244

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationReinforcement Learning

Similar Papers 제목 키워드 기반

Graph Contextual Reinforcement Learning for Efficient Directed Controller Synthesis

2025-12-17 · Toshihide Ubukata, Enhong Mu, Takuto Yamauchi, Mingyue Zhang 외 arxiv

Controller synthesis is a formal method approach for automatically generating Labeled Transition System (LTS) controllers that satisfy specified properties. The efficiency of the synthesis process, however, is critically…

Reinforcement Learning

Exploration Policies for On-the-Fly Controller Synthesis: A Reinforcement Learning Approach

2022-10-07 · Tomás Delgado, Marco Sánchez Sorondo, Víctor Braberman, Sebastián Uchitel

Controller synthesis is in essence a case of model-based planning for non-deterministic environments in which plans (actually ''strategies'') are meant to preserve system goals indefinitely. In the case of supervisory co…

Blockingreinforcement-learningReinforcement Learning (RL)

Bounded Synthesis and Reinforcement Learning of Supervisors for Stochastic Discrete Event Systems with LTL Specifications

2021-05-07 · Ryohei Oura, Toshimitsu Ushio, Ami Sakakibara

In this paper, we consider supervisory control of stochastic discrete event systems (SDESs) under linear temporal logic specifications. Applying the bounded synthesis, we reduce the supervisor synthesis into a problem of…

Formal Language Constrained Markov Decision Processes

2021-01-01 · Eleanor Quint, Dong Xu, Samuel W Flint, Stephen D Scott 외

In order to satisfy safety conditions, an agent may be constrained from acting freely. A safe controller can be designed a priori if an environment is well understood, but not when learning is employed. In particular, re…

MuJoCo

Formal Language Constraints for Markov Decision Processes

2019-10-02 · Eleanor Quint, Dong Xu, Samuel Flint, Stephen Scott 외

In order to satisfy safety conditions, an agent may be constrained from acting freely. A safe controller can be designed a priori if an environment is well understood, but not when learning is employed. In particular, re…

Atari GamesMuJoCo