paper-with-me

홈 › Papers

Regret Minimization in Partially Observable Linear Quadratic Control

2020-01-31 · Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, Anima Anandkumar

We study the problem of regret minimization in partially observable linear quadratic control systems when the model dynamics are unknown a priori. We propose ExpCommit, an explore-then-commit algorithm that learns the model Markov parameters and then follows the principle of optimism in the face of uncertainty to design a controller. We propose a novel way to decompose the regret and provide an end-to-end sublinear regret upper bound for partially observable linear quadratic control. Finally, we provide stability guarantees and establish a regret upper bound of $\tilde{\mathcal{O}}(T^{2/3})$ for ExpCommit, where $T$ is the time horizon of the problem.

📄 PDF Abstract BibTeX arXiv:2002.00082

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Control and Regret Minimization in Linear Quadratic Gaussian (LQG) Setting

2020-03-12 · Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, Anima Anandkumar

We study the problem of adaptive control in partially observable linear quadratic Gaussian control systems, where the model dynamics are unknown a priori. We propose LqgOpt, a novel reinforcement learning algorithm based…

Reinforcement Learning

Logarithmic Regret Bound in Partially Observable Linear Dynamical Systems

2020-03-25 · NeurIPS 2020 12 · Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, Anima Anandkumar

We study the problem of system identification and adaptive control in partially observable linear dynamical systems. Adaptive and closed-loop system identification is a challenging problem due to correlations introduced …

counterfactual

The Partially Observable History Process

2021-11-15 · Dustin Morrill, Amy R. Greenwald, Michael Bowling

We introduce the partially observable history process (POHP) formalism for reinforcement learning. POHP centers around the actions and observations of a single agent and abstracts away the presence of other players witho…

Formreinforcement-learningReinforcement LearningReinforcement Learning (RL)

HSVI can solve zero-sum Partially Observable Stochastic Games

2022-10-26 · Aurélien Delage, Olivier Buffet, Jilles S. Dibangoye, Abdallah Saffidine

State-of-the-art methods for solving 2-player zero-sum imperfect information games rely on linear programming or regret minimization, though not on dynamic programming (DP) or heuristic search (HS), while the latter are …

Decision MakingHeuristic SearchOpen-Ended Question AnsweringSequential Decision Making

Regret Analysis of Learning-Based Linear Quadratic Gaussian Control with Additive Exploration

2023-11-05 · Archith Athrey, Othmane Mazhar, Meichen Guo, Bart De Schutter 외

In this paper, we analyze the regret incurred by a computationally efficient exploration strategy, known as naive exploration, for controlling unknown partially observable systems within the Linear Quadratic Gaussian (LQ…

Efficient Exploration