paper-with-me

Papers

Learning to Control in Metric Space with Optimal Regret

2019-05-05 · Lin F. Yang, Chengzhuo Ni, Mengdi Wang

We study online reinforcement learning for finite-horizon deterministic control systems with {\it arbitrary} state and action spaces. Suppose that the transition dynamics and reward function is unknown, but the state and action space is endowed with a metric that characterizes the proximity between different states and actions. We provide a surprisingly simple upper-confidence reinforcement learning algorithm that uses a function approximation oracle to estimate optimistic Q functions from experiences. We show that the regret of the algorithm after $K$ episodes is $O(HL(KH)^{\frac{d-1}{d}}) $ where $L$ is a smoothness parameter, and $d$ is the doubling dimension of the state-action space with respect to the given metric. We also establish a near-matching regret lower bound. The proposed method can be adapted to work for more structured transition systems, including the finite-state case and the case where value functions are linear combinations of features, where the method also achieve the optimal regret.

📄 PDF Abstract BibTeX arXiv:1905.01576

Code (1)

sjunhongshen/Deterministic-Control-in-Metric-Space

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Regret-optimal Estimation and Control

2021-06-22 · Gautam Goel, Babak Hassibi

We consider estimation and control in linear time-varying dynamical systems from the perspective of regret minimization. Unlike most prior work in this area, we focus on the problem of designing causal estimators and con…

Model Predictive Control

Fast Convergence of Policy Regret in Learning Stochastic Optimal Control

2026-05-25 · Shengbo Wang, Jose Blanchet, Peter Glynn arxiv

Policy learning in modern operations environments faces a fundamental tension between limited operational data and the large, often continuous, state and action spaces over which good decisions must be identified and dep…

Regret-Optimal LQR Control

2021-05-04 · Oron Sabag, Gautam Goel, Sahin Lale, Babak Hassibi

We consider the infinite-horizon LQR control problem. Motivated by competitive analysis in online learning, as a criterion for controller design we introduce the dynamic regret, defined as the difference between the LQR …

Learning Theory

Regret-optimal control in dynamic environments

2020-10-20 · Gautam Goel, Babak Hassibi

We consider control in linear time-varying dynamical systems from the perspective of regret minimization. Unlike most prior work in this area, we focus on the problem of designing an online controller which minimizes reg…

Geometric Exploration for Online Control

2020-10-25 · NeurIPS 2020 12 · Orestis Plevrakis, Elad Hazan

We study the control of an \emph{unknown} linear dynamical system under general convex costs. The objective is minimizing regret vs. the class of disturbance-feedback-controllers, which encompasses all stabilizing linear…