paper-with-me

홈 › Papers

Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies

2026-02-01 · Jin Li, Yue Wu, Mengsha Huang, Yuhao Sun, Hao He, Xianyuan Zhan arxiv

Recurrent neural policies are widely used in partially observable control and meta-RL tasks. Their abilities to maintain internal memory and adapt quickly to unseen scenarios have offered them unparalleled performance when compared to non-recurrent counterparts. However, until today, the underlying mechanisms for their superior generalization and robustness performance remain poorly understood. In this study, by analyzing the hidden state domain of recurrent policies learned over a diverse set of training methods, model architectures, and tasks, we find that stable cyclic structures consistently emerge during interaction with the environment. Such cyclic structures share a remarkable similarity with \textit{limit cycles} in dynamical system analysis, if we consider the policy and the environment as a joint hybrid dynamical system. Moreover, we uncover that the geometry of such limit cycles also has a structured correspondence with the policies' behaviors. These findings offer new perspectives to explain many nice properties of recurrent policies: the emergence of limit cycles stabilizes both the policies' internal memory and the task-relevant environmental states, while suppressing nuisance variability arising from environmental uncertainty; the geometry of limit cycles also encodes relational structures of behaviors, facilitating easier skill adaptation when facing non-stationary environments.

📄 PDF Abstract BibTeX arXiv:2602.01196

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stochastic Physics-Informed Neural Ordinary Differential Equations

2021-09-03 · Jared O'Leary, Joel A. Paulson, Ali Mesbah

Stochastic differential equations (SDEs) are used to describe a wide variety of complex stochastic dynamical systems. Learning the hidden physics within SDEs is crucial for unraveling fundamental understanding of these s…

Neural Co-state Policies: Structuring Hidden States in Recurrent Reinforcement Learning

2026-05-06 · David Leeftink, Max Hinne, Marcel van Gerven arxiv

A key capability of intelligent agents is operating under partial observability: reasoning and acting effectively despite missing or incomplete state observations. While recurrent (memory-based) policies learned via rein…

Reinforcement LearningContinuous Control

Belief-State RWKV for Reinforcement Learning under Partial Observability

2026-04-01 · Liu Xiao arxiv

We propose a stronger formulation of RL on top of RWKV-style recurrent sequence models, in which the fixed-size recurrent state is explicitly interpreted as a belief state rather than an opaque hidden vector. Instead of …

Reinforcement Learning

Recurrent networks, hidden states and beliefs in partially observable environments

2022-08-06 · Gaspard Lambrechts, Adrien Bolland, Damien Ernst

Reinforcement learning aims to learn optimal policies from interaction with environments whose dynamics are unknown. Many methods rely on the approximation of a value function to derive near-optimal policies. In partiall…

Contractive Dynamical Imitation Policies for Efficient Out-of-Sample Recovery

2024-12-10 · Amin Abyaneh, Mahrokh G. Boroujeni, Hsiu-Chin Lin, Giancarlo Ferrari-Trecate

Imitation learning is a data-driven approach to learning policies from expert behavior, but it is prone to unreliable outcomes in out-of-sample (OOS) regions. While previous research relying on stable dynamical systems g…

Imitation Learning