paper-with-me

홈 › Papers

Information maximization for a broad variety of multi-armed bandit games

2025-03-20 · Alex Barbier-Chebbah, Christian L. Vestergaard, Jean-Baptiste Masson

Information and free-energy maximization are physics principles that provide general rules for an agent to optimize actions in line with specific goals and policies. These principles are the building blocks for designing decision-making policies capable of efficient performance with only partial information. Notably, the information maximization principle has shown remarkable success in the classical bandit problem and has recently been shown to yield optimal algorithms for Gaussian and sub-Gaussian reward distributions. This article explores a broad extension of physics-based approaches to more complex and structured bandit problems. To this end, we cover three distinct types of bandit problems, where information maximization is adapted and leads to strong performance. Since the main challenge of information maximization lies in avoiding over-exploration, we highlight how information is tailored at various levels to mitigate this issue, paving the way for more efficient and robust decision-making strategies.

📄 PDF Abstract BibTeX arXiv:2503.15962

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Approximate information maximization for bandit games

2023-10-19 · Alex Barbier-Chebbah, Christian L. Vestergaard, Jean-Baptiste Masson, Etienne Boursier

Entropy maximization and free energy minimization are general physical principles for modeling the dynamics of various physical systems. Notable examples include modeling decision-making within the brain using the free-e…

Decision Making

MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization

2024-12-16 · Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel 외

Reinforcement learning (RL) algorithms aim to balance exploiting the current best strategy with exploring new options that could lead to higher rewards. Most common RL algorithms use undirected exploration, i.e., select …

Multi-Armed BanditsReinforcement Learning (RL)

Horizontally Scalable Submodular Maximization

2016-05-31 · Mario Lucic, Olivier Bachem, Morteza Zadimoghaddam, Andreas Krause

A variety of large-scale machine learning problems can be cast as instances of constrained submodular maximization. Existing approaches for distributed submodular maximization have a critical drawback: The capacity - num…

Trading off rewards and errors in multi-armed bandits

2026-05-01 · Akram Erraqabi, Alessandro Lazaric, Michal Valko, Emma Brunskill 외 arxiv

In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm means accurately and accumulating reward…

Multi-Armed Bandits

No-Regret Learning for Fair Multi-Agent Social Welfare Optimization

2024-05-31 · Mengxiao Zhang, Ramiro Deo-Campo Vuong, Haipeng Luo

We consider the problem of online multi-agent Nash social welfare (NSW) maximization. While previous works of Hossain et al. [2021], Jones et al. [2023] study similar problems in stochastic multi-agent multi-armed bandit…

FairnessMulti-Armed Bandits