paper-with-me

홈 › Papers

A Confirmation of a Conjecture on the Feldman's Two-armed Bandit Problem

2022-06-02 · Zengjing Chen, Yiwei Lin, Jichen Zhang

Myopic strategy is one of the most important strategies when studying bandit problems. In this paper, we consider the two-armed bandit problem proposed by Feldman. With general distributions and utility functions, we obtain a necessary and sufficient condition for the optimality of the myopic strategy. As an application, we could solve Nouiehed and Ross's conjecture for Bernoulli two-armed bandit problems that myopic strategy stochastically maximizes the number of wins.

📄 PDF Abstract BibTeX arXiv:2206.00821

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reducing Dueling Bandits to Cardinal Bandits

2014-05-14 · Nir Ailon, Thorsten Joachims, Zohar Karnin

We present algorithms for reducing the Dueling Bandits problem to the conventional (stochastic) Multi-Armed Bandits problem. The Dueling Bandits problem is an online model of learning with ordinal feedback of the form "A…

Multi-Armed Bandits

A Simple Expression for Mill's Ratio of the Student's $t$-Distribution

2015-02-05 · Francesco Orabona

I show a simple expression of the Mill's ratio of the Student's t-Distribution. I use it to prove Conjecture 1 in P. Auer, N. Cesa-Bianchi, and P. Fischer. Finite-time analysis of the multiarmed bandit problem. Mach. Lea…

Algorithms for multi-armed bandit problems

2014-02-25 · Volodymyr Kuleshov, Doina Precup

Although many algorithms for the multi-armed bandit problem are well-understood theoretically, empirical confirmation of their effectiveness is generally scarce. This paper presents a thorough empirical study of the most…

Multi-Armed Bandits

Beyond the Hazard Rate: More Perturbation Algorithms for Adversarial Multi-armed Bandits

2017-02-17 · Zifan Li, Ambuj Tewari

Recent work on follow the perturbed leader (FTPL) algorithms for the adversarial multi-armed bandit problem has highlighted the role of the hazard rate of the distribution generating the perturbations. Assuming that the …

Multi-Armed Bandits

CM-DQN: A Value-Based Deep Reinforcement Learning Model to Simulate Confirmation Bias

2024-07-10 · Jiacheng Shen, Lihan Feng

In human decision-making tasks, individuals learn through trials and prediction errors. When individuals learn the task, some are more influenced by good outcomes, while others weigh bad outcomes more heavily. Such confi…

Decision MakingDeep Reinforcement Learning