paper-with-me

홈 › Papers

Arm order recognition in multi-armed bandit problem with laser chaos time series

2020-05-26 · Naoki Narisawa, Nicolas Chauvet, Mikio Hasegawa, Makoto Naruse

By exploiting ultrafast and irregular time series generated by lasers with delayed feedback, we have previously demonstrated a scalable algorithm to solve multi-armed bandit (MAB) problems utilizing the time-division multiplexing of laser chaos time series. Although the algorithm detects the arm with the highest reward expectation, the correct recognition of the order of arms in terms of reward expectations is not achievable. Here, we present an algorithm where the degree of exploration is adaptively controlled based on confidence intervals that represent the estimation accuracy of reward expectations. We have demonstrated numerically that our approach did improve arm order recognition accuracy significantly, along with reduced dependence on reward environments, and the total reward is almost maintained compared with conventional MAB methods. This study applies to sectors where the order information is critical, such as efficient allocation of resources in information and communications technology.

📄 PDF Abstract BibTeX arXiv:2005.13085

Code (0)

등록된 구현이 없습니다.

Tasks

Irregular Time SeriesTime SeriesTime Series Analysis

Similar Papers 제목 키워드 기반

Bandit algorithms to emulate human decision making using probabilistic distortions

2016-11-30 · Ravi Kumar Kolla, Prashanth L. A., Aditya Gopalan, Krishna Jagannathan 외

Motivated by models of human decision making proposed to explain commonly observed deviations from conventional expected value preferences, we formulate two stochastic multi-armed bandit problems with distorted probabili…

Decision MakingMulti-Armed Bandits

Indexability of Finite State Restless Multi-Armed Bandit and Rollout Policy

2023-04-30 · Vishesh Mittal, Rahul Meshram, Deepak Dev, Surya Prakash

We consider finite state restless multi-armed bandit problem. The decision maker can act on M bandits out of N bandits in each time step. The play of arm (active arm) yields state dependent rewards based on action and wh…

Parallel photonic accelerator for decision making using optical spatiotemporal chaos

2022-10-12 · Kensei Morijiri, Kento Takehana, Takatomo Mihana, Kazutaka Kanno 외

Photonic accelerators have attracted increasing attention in artificial intelligence applications. The multi-armed bandit problem is a fundamental problem of decision making using reinforcement learning. However, the sca…

Decision Making

Lower Bounds for Multi-armed Bandit with Non-equivalent Multiple Plays

2015-07-17 · Aleksandr Vorobev, Gleb Gusev

We study the stochastic multi-armed bandit problem with non-equivalent multiple plays where, at each step, an agent chooses not only a set of arms, but also their order, which influences reward distribution. In several p…

Lifelong Learning in Multi-Armed Bandits

2020-12-28 · Matthieu Jedor, Jonathan Louëdec, Vianney Perchet

Continuously learning and leveraging the knowledge accumulated from prior tasks in order to improve future performance is a long standing machine learning problem. In this paper, we study the problem in the multi-armed b…

Lifelong learningMulti-Armed Bandits