paper-with-me

홈 › Papers

Learning with Bandit Feedback in Potential Games

2017-12-01 · NeurIPS 2017 12 · Amélie Heliou, Johanne Cohen, Panayotis Mertikopoulos

This paper examines the equilibrium convergence properties of no-regret learning with exponential weights in potential games. To establish convergence with minimal information requirements on the players' side, we focus on two frameworks: the semi-bandit case (where players have access to a noisy estimate of their payoff vectors, including strategies they did not play), and the bandit case (where players are only able to observe their in-game, realized payoffs). In the semi-bandit case, we show that the induced sequence of play converges almost surely to a Nash equilibrium at a quasi-exponential rate. In the bandit case, the same result holds for approximate Nash equilibria if we introduce a constant exploration factor that guarantees that action choice probabilities never become arbitrarily small. In particular, if the algorithm is run with a suitably decreasing exploration factor, the sequence of play converges to a bona fide Nash equilibrium with probability 1.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Randomised Optimism via Competitive Co-Evolution for Matrix Games with Bandit Feedback

2025-05-19 · Shishen Lin

Learning in games is a fundamental problem in machine learning and artificial intelligence, with numerous applications~\citep{silver2016mastering,schrittwieser2020mastering}. This work investigates two-player zero-sum ma…

Evolutionary Algorithms

Adversarial Learning in Games with Bandit Feedback: Logarithmic Pure-Strategy Maximin Regret

2026-02-06 · Shinji Ito, Haipeng Luo, Arnab Maiti, Taira Tsuchiya 외 arxiv

Learning to play zero-sum games is a fundamental problem in game theory and machine learning. While significant progress has been made in minimizing external regret in the self-play settings or with full-information feed…

Offline congestion games: How feedback type affects data coverage requirement

2022-10-24 · Haozhe Jiang, Qiwen Cui, Zhihan Xiong, Maryam Fazel 외

This paper investigates when one can efficiently recover an approximate Nash Equilibrium (NE) in offline congestion games. The existing dataset coverage assumption in offline general-sum games inevitably incurs a depende…

Vocal Bursts Type Prediction

No-regret learning for repeated non-cooperative games with lossy bandits

2022-05-14 · Wenting Liu, Jinlong Lei, Peng Yi, Yiguang Hong

This paper considers no-regret learning for repeated continuous-kernel games with lossy bandit feedback. Since it is difficult to give the explicit model of the utility functions in dynamic environments, the players' act…

Management

Sample-Efficient Learning of Stackelberg Equilibria in General-Sum Games

2021-02-23 · NeurIPS 2021 12 · Yu Bai, Chi Jin, Huan Wang, Caiming Xiong

Real world applications such as economics and policy making often involve solving multi-agent games with two unique features: (1) The agents are inherently asymmetric and partitioned into leaders and followers; (2) The a…