paper-with-me

홈 › Papers

Approximate information maximization for bandit games

2023-10-19 · Alex Barbier-Chebbah, Christian L. Vestergaard, Jean-Baptiste Masson, Etienne Boursier

Entropy maximization and free energy minimization are general physical principles for modeling the dynamics of various physical systems. Notable examples include modeling decision-making within the brain using the free-energy principle, optimizing the accuracy-complexity trade-off when accessing hidden variables with the information bottleneck principle (Tishby et al., 2000), and navigation in random environments using information maximization (Vergassola et al., 2007). Built on this principle, we propose a new class of bandit algorithms that maximize an approximation to the information of a key variable within the system. To this end, we develop an approximated analytical physics-based representation of an entropy to forecast the information gain of each action and greedily choose the one with the largest information gain. This method yields strong performances in classical bandit settings. Motivated by its empirical success, we prove its asymptotic optimality for the two-armed bandit problem with Gaussian rewards. Owing to its ability to encompass the system's properties in a global physical functional, this approach can be efficiently adapted to more complex bandit settings, calling for further investigation of information maximization approaches for multi-armed bandit problems.

📄 PDF Abstract BibTeX arXiv:2310.12563

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Information maximization for a broad variety of multi-armed bandit games

2025-03-20 · Alex Barbier-Chebbah, Christian L. Vestergaard, Jean-Baptiste Masson

Information and free-energy maximization are physics principles that provide general rules for an agent to optimize actions in line with specific goals and policies. These principles are the building blocks for designing…

Decision Making

Offline congestion games: How feedback type affects data coverage requirement

2022-10-24 · Haozhe Jiang, Qiwen Cui, Zhihan Xiong, Maryam Fazel 외

This paper investigates when one can efficiently recover an approximate Nash Equilibrium (NE) in offline congestion games. The existing dataset coverage assumption in offline general-sum games inevitably incurs a depende…

Vocal Bursts Type Prediction

Learning with Bandit Feedback in Potential Games

2017-12-01 · NeurIPS 2017 12 · Amélie Heliou, Johanne Cohen, Panayotis Mertikopoulos

This paper examines the equilibrium convergence properties of no-regret learning with exponential weights in potential games. To establish convergence with minimal information requirements on the players' side, we focus …

Sample-Efficient Learning of Correlated Equilibria in Extensive-Form Games

2022-05-15 · Ziang Song, Song Mei, Yu Bai

Imperfect-Information Extensive-Form Games (IIEFGs) is a prevalent model for real-world games involving imperfect information and sequential plays. The Extensive-Form Correlated Equilibrium (EFCE) has been proposed as a …

Form

Protocols for Verifying Smooth Strategies in Bandits and Games

2025-07-08 · Miranda Christ, Daniel Reichman, Jonathan Shafer arxiv

We study protocols for verifying approximate optimality of strategies in multi-armed bandits and normal-form games. As the number of actions available to each player is often large, we seek protocols where the number of …

Multi-Armed Bandits