paper-with-me

홈 › Papers

Taming "data-hungry" reinforcement learning? Stability in continuous state-action spaces

2024-01-10 · Yaqi Duan, Martin J. Wainwright

We introduce a novel framework for analyzing reinforcement learning (RL) in continuous state-action spaces, and use it to prove fast rates of convergence in both off-line and on-line settings. Our analysis highlights two key stability properties, relating to how changes in value functions and/or policies affect the Bellman operator and occupation measures. We argue that these properties are satisfied in many continuous state-action Markov decision processes, and demonstrate how they arise naturally when using linear function approximation methods. Our analysis offers fresh perspectives on the roles of pessimism and optimism in off-line and on-line RL, and highlights the connection between off-line RL and transfer learning.

📄 PDF Abstract BibTeX arXiv:2401.05233

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning (RL)Transfer Learning

Similar Papers 제목 키워드 기반

RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin Learning

2026-07-21 · Yiwei Zhou, Ziheng Chen arxiv

We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift. A threshold determines where the taming turns on, wh…

From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight

2026-03-15 · Xiaoliang Fu, Jiaye Lin, Yangyi Fang, Chaowen Hu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed a leap in Large Language Model (LLM) reasoning, yet its optimization dynamics remain fragile. Standard algorithms like GRPO enforce stability via "hard …

Reinforcement Learning

On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator

2019-01-11 · Qi Cai, Mingyi Hong, Yongxin Chen, Zhaoran Wang

We study the global convergence of generative adversarial imitation learning for linear quadratic regulators, which is posed as minimax optimization. To address the challenges arising from non-convex-concave geometry, we…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Taming Lagrangian Chaos with Multi-Objective Reinforcement Learning

2022-12-19 · Chiara Calascibetta, Luca Biferale, Francesco Borra, Antonio Celani 외

We consider the problem of two active particles in 2D complex flows with the multi-objective goals of minimizing both the dispersion rate and the energy consumption of the pair. We approach the problem by means of Multi …

Multi-Objective Reinforcement LearningQ-Learningreinforcement-learningReinforcement Learning+1

Taming Continuous Posteriors for Latent Variational Dialogue Policies

2022-05-16 · Marin Vlastelica, Patrick Ernst, György Szarvas

Utilizing amortized variational inference for latent-action reinforcement learning (RL) has been shown to be an effective approach in Task-oriented Dialogue (ToD) systems for optimizing dialogue success. Until now, categ…

reinforcement-learningReinforcement Learning (RL)Variational Inference