paper-with-me

홈 › Papers

Near-continuous time Reinforcement Learning for continuous state-action spaces

2023-09-06 · Lorenzo Croissant, Marc Abeille, Bruno Bouchard

We consider the Reinforcement Learning problem of controlling an unknown dynamical system to maximise the long-term average reward along a single trajectory. Most of the literature considers system interactions that occur in discrete time and discrete state-action spaces. Although this standpoint is suitable for games, it is often inadequate for mechanical or digital systems in which interactions occur at a high frequency, if not in continuous time, and whose state spaces are large if not inherently continuous. Perhaps the only exception is the Linear Quadratic framework for which results exist both in discrete and continuous time. However, its ability to handle continuous states comes with the drawback of a rigid dynamic and reward structure. This work aims to overcome these shortcomings by modelling interaction times with a Poisson clock of frequency $\varepsilon^{-1}$, which captures arbitrary time scales: from discrete ($\varepsilon=1$) to continuous time ($\varepsilon\downarrow0$). In addition, we consider a generic reward function and model the state dynamics according to a jump process with an arbitrary transition kernel on $\mathbb{R}^d$. We show that the celebrated optimism protocol applies when the sub-tasks (learning and planning) can be performed effectively. We tackle learning within the eluder dimension framework and propose an approximate planning method based on a diffusive limit approximation of the jump process. Overall, our algorithm enjoys a regret of order $\tilde{\mathcal{O}}(\varepsilon^{1/2} T+\sqrt{T})$. As the frequency of interactions blows up, the approximation error $\varepsilon^{1/2} T$ vanishes, showing that $\tilde{\mathcal{O}}(\sqrt{T})$ is attainable in near-continuous time.

📄 PDF Abstract BibTeX arXiv:2309.02815

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

Regret Analysis of Certainty Equivalence Policies in Continuous-Time Linear-Quadratic Systems

2022-06-09 · Mohamad Kazem Shirani Faradonbeh

This work theoretically studies a ubiquitous reinforcement learning policy for controlling the canonical model of continuous-time stochastic linear-quadratic systems. We show that randomized certainty equivalent policy a…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon

2020-06-27 · Matteo Basei, Xin Guo, Anran Hu, Yufei Zhang

We study finite-time horizon continuous-time linear-quadratic reinforcement learning problems in an episodic setting, where both the state and control coefficients are unknown to the controller. We first propose a least-…

parameter estimationReinforcement Learning (RL)

Efficient Exploration in Continuous-time Model-based Reinforcement Learning

2023-10-30 · NeurIPS 2023 11

Reinforcement learning algorithms typically consider discrete-time dynamics, even though the underlying systems are often continuous in time. In this paper, we introduce a model-based reinforcement learning algorithm tha…

Efficient ExplorationGaussian ProcessesModel-based Reinforcement Learningreinforcement-learning+1

Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control

2025-10-20 · Chengxiu Hua, Jiawen Gu, Yushun Tang arxiv

Reinforcement learning (RL) has achieved significant success across a wide range of domains, however, most existing methods are formulated in discrete time. In this work, we introduce a novel RL method for continuous-tim…

Reinforcement Learning

Policy Transfer for Continuous-Time Reinforcement Learning: A (Rough) Differential Equation Approach

2025-10-16 · Xin Guo, Zijiu Lyu arxiv

This paper studies policy transfer, one of the well-known transfer learning techniques adopted in large language models, for continuous-time reinforcement learning problems. In the case of continuous-time linear-quadrati…

Reinforcement LearningTransfer Learning