paper-with-me

홈 › Papers

Bandit Linear Control

2020-07-01 · NeurIPS 2020 12 · Asaf Cassel, Tomer Koren

We consider the problem of controlling a known linear dynamical system under stochastic noise, adversarially chosen costs, and bandit feedback. Unlike the full feedback setting where the entire cost function is revealed after each decision, here only the cost incurred by the learner is observed. We present a new and efficient algorithm that, for strongly convex and smooth costs, obtains regret that grows with the square root of the time horizon $T$. We also give extensions of this result to general convex, possibly non-smooth costs, and to non-stochastic system noise. A key component of our algorithm is a new technique for addressing bandit optimization of loss functions with memory.

📄 PDF Abstract BibTeX arXiv:2007.00759

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimal Rates for Bandit Nonstochastic Control

2023-05-24 · NeurIPS 2023 11

Linear Quadratic Regulator (LQR) and Linear Quadratic Gaussian (LQG) control are foundational and extensively researched problems in optimal control. We investigate LQR and LQG problems with semi-adversarial perturbation…

Online Limited Memory Neural-Linear Bandits

2021-01-01 · Tom Zahavy, Ofir Nabati, Leor Cohen, Shie Mannor

We study neural-linear bandits for solving problems where both exploration and representation learning play an important role. Neural-linear bandits leverage the representation power of deep neural networks and combine i…

Efficient ExplorationMulti-Armed BanditsRepresentation LearningSentiment Analysis

Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits

2020-10-08 · Yu-Heng Hung, Ping-Chun Hsieh, Xi Liu, P. R. Kumar

Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit trade-off in linear bandits problems as wel…

Computational Efficiency

When Are Linear Stochastic Bandits Attackable?

2021-10-18 · Huazheng Wang, Haifeng Xu, Hongning Wang

We study adversarial attacks on linear stochastic bandits: by manipulating the rewards, an adversary aims to control the behaviour of the bandit algorithm. Perhaps surprisingly, we first show that some attack goals can n…

Decision MakingRecommendation Systems

Non-Stochastic Control with Bandit Feedback

2020-08-12 · NeurIPS 2020 12 · Paula Gradu, John Hallman, Elad Hazan

We study the problem of controlling a linear dynamical system with adversarial perturbations where the only feedback available to the controller is the scalar loss, and the loss function itself is unknown. For this probl…