paper-with-me

홈 › Papers

Non-Stochastic Control with Bandit Feedback

2020-08-12 · NeurIPS 2020 12 · Paula Gradu, John Hallman, Elad Hazan

We study the problem of controlling a linear dynamical system with adversarial perturbations where the only feedback available to the controller is the scalar loss, and the loss function itself is unknown. For this problem, with either a known or unknown system, we give an efficient sublinear regret algorithm. The main algorithmic difficulty is the dependence of the loss on past controls. To overcome this issue, we propose an efficient algorithm for the general setting of bandit convex optimization for loss functions with memory, which may be of independent interest.

📄 PDF Abstract BibTeX arXiv:2008.05523

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bandit Linear Control

2020-07-01 · NeurIPS 2020 12 · Asaf Cassel, Tomer Koren

We consider the problem of controlling a known linear dynamical system under stochastic noise, adversarially chosen costs, and bandit feedback. Unlike the full feedback setting where the entire cost function is revealed …

Bandit Structured Prediction for Neural Sequence-to-Sequence Learning

2017-04-21 · ACL 2017 7 · Julia Kreutzer, Artem Sokolov, Stefan Riezler

Bandit structured prediction describes a stochastic optimization framework where learning is performed from partial feedback. This feedback is received in the form of a task loss evaluation to a predicted output structur…

Domain AdaptationMachine TranslationStochastic OptimizationStructured Prediction+1

Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback

2023-03-23 · Mohammad Pedramfar, Vaneet Aggarwal

This paper investigates the problem of combinatorial multiarmed bandits with stochastic submodular (in expectation) rewards and full-bandit delayed feedback, where the delayed feedback is assumed to be composite and anon…

A Regret Perspective on Online Selective Generation

2025-06-16 · Minjae Lee, Yoonjae Jung, Sangdon Park

Large language generative models increasingly interact with humans, while their falsified responses raise concerns. To address this hallucination effect, selectively abstaining from answering, called selective generation…

HallucinationLEMMA

Nearly Tight Bounds for Cross-Learning Contextual Bandits with Graphical Feedback

2025-02-07 · Ruiyuan Huang, Zengfeng Huang

The cross-learning contextual bandit problem with graphical feedback has recently attracted significant attention. In this setting, there is a contextual bandit with a feedback graph over the arms, and pulling an arm rev…

Multi-Armed Bandits