Functional Bandits
We introduce the functional bandit problem, where the objective is to find an arm that optimises a known functional of the unknown arm-reward distributions. These problems arise in many settings such as maximum entropy methods in natural language processing, and risk-averse decision-making, but current best-arm identification techniques fail in these domains. We propose a new approach, that combines functional estimation and arm elimination, to tackle this problem. This method achieves provably efficient performance guarantees. In addition, we illustrate this method on a number of important functionals in risk management and information theory, and refine our generic theoretical results in those cases.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingManagementSimilar Papers 제목 키워드 기반
Efficient Contextual Bandits with Continuous Actions
We create a computationally tractable algorithm for contextual bandits with continuous actions having unknown structure. Our reduction-style algorithm composes with most supervised learning representations. We prove that…
Multi-Armed BanditsSignature Approach for Contextual Bandits with Nonlinear and Path-dependent Rewards
We study contextual bandits with nonlinear and path-dependent rewards through a novel signature-transform-based approach. Leveraging the universal nonlinearity property of signatures, we approximate continuous path-depen…
Off-Policy Risk Assessment in Contextual Bandits
Even when unable to run experiments, practitioners can evaluate prospective policies, using previously logged data. However, while the bandits literature has adopted a diverse set of objectives, most research on off-poli…
Multi-Armed BanditsOff-policy evaluationContextual Online Decision Making with Infinite-Dimensional Functional Regression
Contextual sequential decision-making problems play a crucial role in machine learning, encompassing a wide range of downstream applications such as bandits, sequential hypothesis testing and online risk control. These a…
Decision MakingMulti-Armed BanditsregressionSequential Decision MakingA Deep Bayesian Bandits Approach for Anticancer Therapy: Exploration via Functional Prior
Learning personalized cancer treatment with machine learning holds great promise to improve cancer patients' chance of survival. Despite recent advances in machine learning and precision oncology, this approach remains c…
BIG-bench Machine LearningDrug Response Prediction