paper-with-me

Papers

Outbidding and Outbluffing Elite Humans: Mastering Liar's Poker via Self-Play and Reinforcement Learning

2025-11-05 · Richard Dewey, Janos Botyanszki, Ciamac C. Moallemi, Andrew T. Zheng arxiv

AI researchers have long focused on poker-like games as a testbed for environments characterized by multi-player dynamics, imperfect information, and reasoning under uncertainty. While recent breakthroughs have matched elite human play at no-limit Texas hold'em, the multi-player dynamics are subdued: most hands converge quickly with only two players engaged through multiple rounds of bidding. In this paper, we present Solly, the first AI agent to achieve elite human play in reduced-format Liar's Poker, a game characterized by extensive multi-player engagement. We trained Solly using self-play with a model-free, actor-critic, deep reinforcement learning algorithm. Solly played at an elite human level as measured by win rate (won over 50% of hands) and equity (money won) in heads-up and multi-player Liar's Poker. Solly also outperformed large language models (LLMs), including those with reasoning abilities, on the same metrics. Solly developed novel bidding strategies, randomized play effectively, and was not easily exploitable by world-class human players.

📄 PDF Abstract BibTeX arXiv:2511.03724

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

End-to-end Music Remastering System Using Self-supervised and Adversarial Training

2022-02-17 · Junghyun Koo, Seungryeol Paik, Kyogu Lee

Mastering is an essential step in music production, but it is also a challenging task that has to go through the hands of experienced audio engineers, where they adjust tone, space, and volume of a song. Remastering foll…

A Comparison Between Decision Trees and Decision Tree Forest Models for Software Development Effort Estimation

2015-08-28 · Ali Bou Nassif, Mohammad Azzeh, Luiz Fernando Capretz, Danny Ho

Accurate software effort estimation has been a challenge for many software practitioners and project managers. Underestimation leads to disruption in the projects estimated cost and delivery. On the other hand, overestim…

regression

Discovering the Elite Hypervolume by Leveraging Interspecies Correlation

2018-04-11 · Vassilis Vassiliades, Jean-Baptiste Mouret

Evolution has produced an astonishing diversity of species, each filling a different niche. Algorithms like MAP-Elites mimic this divergent evolutionary process to find a set of behaviorally diverse but high-performing s…

Diversity

Multi-Task Multi-Behavior MAP-Elites

2023-05-02 · Anne, Mouret

We propose Multi-Task Multi-Behavior MAP-Elites, a variant of MAP-Elites that finds a large number of high-quality solutions for a large set of tasks (optimization problems from a given family). It combines the original …

Diversity

Self-Referential Quality Diversity Through Differential Map-Elites

2021-07-11 · Tae Jong Choi, Julian Togelius

Differential MAP-Elites is a novel algorithm that combines the illumination capacity of CVT-MAP-Elites with the continuous-space optimization capacity of Differential Evolution. The algorithm is motivated by observations…

Diversity