paper-with-me

Papers

Continuous-time reinforcement learning: ellipticity enables model-free value function approximation

2026-02-06 · Wenlong Mou arxiv

We study off-policy reinforcement learning for controlling continuous-time Markov diffusion processes with discrete-time observations and actions. We consider model-free algorithms with function approximation that learn value and advantage functions directly from data, without unrealistic structural assumptions on the dynamics. Leveraging the ellipticity of the diffusions, we establish a new class of Hilbert-space positive definiteness and boundedness properties for the Bellman operators. Based on these properties, we propose the Sobolev-prox fitted $q$-learning algorithm, which learns value and advantage functions by iteratively solving least-squares regression problems. We derive oracle inequalities for the estimation error, governed by (i) the best approximation error of the function classes, (ii) their localized complexity, (iii) exponentially decaying optimization error, and (iv) numerical discretization error. These results identify ellipticity as a key structural property that renders reinforcement learning with function approximation for Markov diffusions no harder than supervised learning.

📄 PDF Abstract BibTeX arXiv:2602.06930

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Statistical guarantees for continuous-time policy evaluation: blessing of ellipticity and new tradeoffs

2025-02-06 · Wenlong Mou

We study the estimation of the value function for continuous-time Markov diffusion processes using a single, discretely observed ergodic trajectory. Our work provides non-asymptotic statistical guarantees for the least-s…

Deep Q-Learning on Hölder Spaces

2026-06-15 · Qian Qi arxiv

We study the operator-theoretic core of Q-learning in continuous-time stochastic control with continuous states and actions. In value-based reinforcement learning, each Q-learning or DQN update is built from a Bellman op…

Reinforcement Learning

Towards solving model bias in cosmic shear forward modeling

2022-10-28 · Benjamin Remy, Francois Lanusse, Jean-Luc Starck

As the volume and quality of modern galaxy surveys increase, so does the difficulty of measuring the cosmological signal imprinted in galaxy shapes. Weak gravitational lensing sourced by the most massive structures in th…

Deterministic Policy Gradient for Reinforcement Learning with Continuous Time and State

2025-09-28 · Ziheng Cheng, Xin Guo, Yufei Zhang arxiv

The theory of continuous-time reinforcement learning (RL) has progressed rapidly in recent years. While the ultimate objective of RL is typically to learn deterministic control policies, most existing continuous-time RL …

General Reinforcement Learning

Formal Controller Synthesis for Continuous-Space MDPs via Model-Free Reinforcement Learning

2020-03-02 · Abolfazl Lavaei, Fabio Somenzi, Sadegh Soudjani, Ashutosh Trivedi 외

A novel reinforcement learning scheme to synthesize policies for continuous-space Markov decision processes (MDPs) is proposed. This scheme enables one to apply model-free, off-the-shelf reinforcement learning algorithms…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)