paper-with-me

Papers

Regret-Optimal Control under Partial Observability

2023-11-10 · Joudi Hajar, Oron Sabag, Babak Hassibi

This paper studies online solutions for regret-optimal control in partially observable systems over an infinite-horizon. Regret-optimal control aims to minimize the difference in LQR cost between causal and non-causal controllers while considering the worst-case regret across all $\ell_2$-norm-bounded disturbance and measurement sequences. Building on ideas from Sabag et al., 2023, on the the full-information setting, our work extends the framework to the scenario of partial observability (measurement-feedback). We derive an explicit state-space solution when the non-causal solution is the one that minimizes the $\mathcal H_2$ criterion, and demonstrate its practical utility on several practical examples. These results underscore the framework's significant relevance and applicability in real-world systems.

📄 PDF Abstract BibTeX arXiv:2311.06433

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Bandits to Experts: A Tale of Domination and Independence

2013-07-17 · NeurIPS 2013 12 · Noga Alon, Nicolò Cesa-Bianchi, Claudio Gentile, Yishay Mansour

We consider the partial observability model for multi-armed bandits, introduced by Mannor and Shamir. Our main result is a characterization of regret in the directed observability model in terms of the dominating and ind…

Multi-Armed Bandits

What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty

2026-03-03 · Aran Nayebi arxiv

As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? Classical results show that optimal control can be implemented using belief states…

Efficient RL with Impaired Observability: Learning to Act with Delayed and Missing State Observations

2023-09-21 · NeurIPS 2023 11

In real-world reinforcement learning (RL) systems, various forms of {\it impaired observability} can complicate matters. These situations arise when an agent is unable to observe the most recent state of the system due t…

Online learning with noisy side observations

2026-04-15 · Tomáš Kocák, Gergely Neu, Michal Valko arxiv

We propose a new partial-observability model for online learning problems where the learner, besides its own loss, also observes some noisy feedback about the other actions, depending on the underlying structure of the p…

Efficient and Optimal No-Regret Caching under Partial Observation

2025-03-04 · Younes Ben Mazziane, Francescomaria Faticanti, Sara Alouf, Giovanni Neglia

Online learning algorithms have been successfully used to design caching policies with sublinear regret in the total number of requests, with no statistical assumption about the request sequence. Most existing algorithms…