paper-with-me

Papers

Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation

2025-08-04 · Runze Zhao, Yue Yu, Ruhan Wang, Chunfeng Huang, Dongruo Zhou arxiv

Continuous-time reinforcement learning (CTRL) provides a natural framework for sequential decision-making in dynamic environments where interactions evolve continuously over time. While CTRL has shown growing empirical success, its ability to adapt to varying levels of problem difficulty remains poorly understood. In this work, we investigate the instance-dependent behavior of CTRL and introduce a simple, model-based algorithm built on maximum likelihood estimation (MLE) with a general function approximator. Unlike existing approaches that estimate system dynamics directly, our method estimates the state marginal density to guide learning. We establish instance-dependent performance guarantees by deriving a regret bound that scales with the total reward variance and measurement resolution. Notably, the regret becomes independent of the specific measurement strategy when the observation frequency adapts appropriately to the problem's complexity. To further improve performance, our algorithm incorporates a randomized measurement schedule that enhances sample efficiency without increasing measurement cost. These results highlight a new direction for designing CTRL algorithms that automatically adjust their learning behavior based on the underlying difficulty of the environment.

📄 PDF Abstract BibTeX arXiv:2508.02103

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Maximum diffusion reinforcement learning

2023-09-26 · Thomas A. Berrueta, Allison Pinosky, Todd D. Murphey

Robots and animals both experience the world through their bodies and senses. Their embodiment constrains their experiences, ensuring they unfold continuously in space and time. As a result, the experiences of embodied a…

Decision Makingreinforcement-learningReinforcement LearningSelf-Driving Cars

Logarithmic regret bounds for continuous-time average-reward Markov decision processes

2022-05-23 · Xuefeng Gao, Xun Yu Zhou

We consider reinforcement learning for continuous-time Markov decision processes (MDPs) in the infinite-horizon, average-reward setting. In contrast to discrete-time MDPs, a continuous-time process moves to a state and s…

Point Processesreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Deep Multi-Agent Reinforcement Learning with Hybrid Action Spaces based on Maximum Entropy

2022-06-10 · Hongzhi Hua, Kaigui Wu, Guixuan Wen

Multi-agent deep reinforcement learning has been applied to address a variety of complex problems with either discrete or continuous action spaces and achieved great success. However, most real-world environments cannot …

Deep Reinforcement LearningMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+1

Near Instance-Optimal PAC Reinforcement Learning for Deterministic MDPs

2022-03-17 · Andrea Tirinzoni, Aymen Al-Marjani, Emilie Kaufmann

In probably approximately correct (PAC) reinforcement learning (RL), an agent is required to identify an $\epsilon$-optimal policy with probability $1-\delta$. While minimax optimal algorithms exist for this problem, its…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

SCC-rFMQ Learning in Cooperative Markov Games with Continuous Actions

2018-09-18 · Chengwei Zhang, Xiaohong Li, Jianye Hao, Siqi Chen 외

Although many reinforcement learning methods have been proposed for learning the optimal solutions in single-agent continuous-action domains, multiagent coordination domains with continuous actions have received relative…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)