paper-with-me

홈 › Papers

Dynamic Latent Routing

2026-05-14 · Fangyuan Yu, Xin Su, Amir Abdullah arxiv

We investigate the temporal concatenation of sub-policies in Markov Decision Processes (MDP) with time-varying reward functions. We introduce General Dijkstra Search (GDS), and prove that globally optimal goal-reaching policies can be recovered through temporal composition of intermediate optimal sub-policies. Motivated by the "search, select, update" principle underlying GDS, we propose Dynamic Latent Routing (DLR), a language-model post-training method that jointly learns discrete latent codes, routing policies, and model parameters through dynamic search in a single training stage. In low-data fine-tuning settings, DLR matches or outperforms supervised fine-tuning across four datasets and six models, achieving a mean gain of +6.6 percentage points, while prior discrete-latent baselines consistently underperform SFT. Mechanistic analyses and targeted code ablations show that DLR learns structured routing behaviors with distinct causal roles.

📄 PDF Abstract BibTeX arXiv:2605.14323

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariate Time Series Forecasting

2026-08-21 · Lan Guo, Jie Xiao, Zhao Su, Jun Shen 외 arxiv

In non-stationary multivariate time series, different variables and samples often exhibit heterogeneous latent dynamic states, while existing deep forecasting models usually compress them into a unified end-to-end mappin…

Multivariate Time Series Forecasting

LAR-MoE: Latent-Aligned Routing for Mixture of Experts in Robotic Imitation Learning

2026-03-09 · Ariel Rodriguez, Chenpan Li, Lorenzo Mazza, Rayan Younis 외 arxiv

Imitation learning enables robots to acquire manipulation skills from demonstrations, yet deploying a policy across tasks with heterogeneous dynamics remains challenging, as models tend to average over distinct behaviora…

ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces

2026-02-12 · Xin Xu, Tong Yu, Xiang Chen, Haoliang Wang 외 arxiv

Recent work explores latent reasoning to improve reasoning efficiency by replacing explicit reasoning trajectories with continuous representations in a latent space, yet its effectiveness varies across settings. Analysis…

TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization

2026-06-04 · Liting Zhang, Shiwan Zhao, Xuyang Zhao, Zichen Xu 외 arxiv

Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive reasoning by operating over continuous representations. However, the in…

Reinforcement Learning

Statistic-Augmented, Decoupled MoE Routing and Aggregating in Autonomous Driving

2025-12-07 · Wei-Bin Kou, Guangxu Zhu, Jingreng Lei, Chen Zhang 외 arxiv

Autonomous driving (AD) scenarios are inherently complex and diverse, posing significant challenges for a single deep learning model to effectively cover all possible conditions, such as varying weather, traffic densitie…

Semantic SegmentationAutonomous Driving