paper-with-me

Papers

Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits

2020-10-08 · Yu-Heng Hung, Ping-Chun Hsieh, Xi Liu, P. R. Kumar

Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit trade-off in linear bandits problems as well as generalized linear bandits problems. We develop novel index policies that we prove achieve order-optimality, and show that they achieve empirical performance competitive with the state-of-the-art benchmark methods in extensive experiments. The new policies achieve this with low computation time per pull for linear bandits, and thereby resulting in both favorable regret as well as computational efficiency.

📄 PDF Abstract BibTeX arXiv:2010.04091

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Reward-Biased Maximum Likelihood Estimation for Neural Contextual Bandits

2022-03-08 · Yu-Heng Hung, Ping-Chun Hsieh

Reward-biased maximum likelihood estimation (RBMLE) is a classic principle in the adaptive control literature for tackling explore-exploit trade-offs. This paper studies the stochastic contextual bandit problem with gene…

Multi-Armed Bandits

Reward Biased Maximum Likelihood Estimation for Reinforcement Learning

2020-11-16 · Akshay Mete, Rahul Singh, Xi Liu, P. R. Kumar

The Reward-Biased Maximum Likelihood Estimate (RBMLE) for adaptive control of Markov chains was proposed to overcome the central obstacle of what is variously called the fundamental "closed-identifiability problem" of ad…

Multi-Armed Banditsreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

A new adjusted maximum likelihood method for the Fay–Herriot small area model

2013-11-20 · Journal of Multivariate Analysis 2013 11 · Masayo Yoshimori ∗, Partha Lahiri

In the context of the Fay–Herriot model, a mixed regression model routinely used to combine information from various sources in small area estimation, certain adjustments to a standard likelihood (e.g., profile, residu…

regression

Exploration Through Bias: Revisiting Biased Maximum Likelihood Estimation in Stochastic Multi-Armed Bandits

2020-01-01 · ICML 2020 1 · Xi Liu, Ping-Chun Hsieh, Yu Heng Hung, Anirban Bhattacharya 외

We propose a new family of bandit algorithms, that are formulated in a general way based on the Biased Maximum Likelihood Estimation (BMLE) method originally appearing in the adaptive control literature. We design the re…

Multi-Armed Bandits

Exploration Through Reward Biasing: Reward-Biased Maximum Likelihood Estimation for Stochastic Multi-Armed Bandits

2019-07-02 · Xi Liu, Ping-Chun Hsieh, Anirban Bhattacharya, P. R. Kumar

Inspired by the Reward-Biased Maximum Likelihood Estimate method of adaptive control, we propose RBMLE -- a novel family of learning algorithms for stochastic multi-armed bandits (SMABs). For a broad range of SMABs inclu…

Multi-Armed Bandits