paper-with-me

홈 › Papers

Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures

2025-06-07 · Clément Hongler, Andrew Emil

Large Language Models (LLMs) define probability measures on text. By considering the implicit knowledge question of what it means for an LLM to know such a measure and what it entails algorithmically, we are naturally led to formulate a series of tasks that go beyond generative sampling, involving forms of summarization, counterfactual thinking, anomaly detection, originality search, reverse prompting, debating, creative solving, etc. These tasks can be formulated as games based on LLM measures, which we call Cross-Entropy (Xent) Games. Xent Games can be single-player or multi-player. They involve cross-entropy scores and cross-entropy constraints, and can be expressed as simple computational graphs and programs. We show the Xent Game space is large enough to contain a wealth of interesting examples, while being constructible from basic game-theoretic consistency axioms. We then discuss how the Xent Game space can be used to measure the abilities of LLMs. This leads to the construction of Xent Game measures: finite families of Xent Games that can be used as capability benchmarks, built from a given scope, by extracting a covering measure. To address the unbounded scope problem associated with the challenge of measuring general abilities, we propose to explore the space of Xent Games in a coherent fashion, using ideas inspired by evolutionary dynamics.

📄 PDF Abstract BibTeX arXiv:2506.06832

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly Detectioncounterfactual

Similar Papers 제목 키워드 기반

Policy Optimization in Zero-Sum Markov Games: Fictitious Self-Play Provably Attains Nash Equilibria

2021-01-01 · Boyi Liu, Zhuoran Yang, Zhaoran Wang

Fictitious Self-Play (FSP) has achieved significant empirical success in solving extensive-form games. However, from a theoretical perspective, it remains unknown whether FSP is guaranteed to converge to Nash equilibria…

Uncoupled and Convergent Learning in Two-Player Zero-Sum Markov Games with Bandit Feedback

2023-03-05 · NeurIPS 2023 11

We revisit the problem of learning in two-player zero-sum Markov games, focusing on developing an algorithm that is uncoupled, convergent, and rational, with non-asymptotic convergence rates. We start from the case of st…

Cognitive Training for Language Models: Towards General Capabilities via Cross-Entropy Games

2026-03-23 · Clément Hongler, Franck Gabriel, Valentin Hartmann, Arthur Renard 외 arxiv

Defining a constructive process to build general capabilities for language models in an automatic manner is considered an open problem in artificial intelligence. Towards this, we consider the problem of building a curri…

Entropy Non-increasing Games for the Improvement of Dataflow Programming

2017-02-14 · Norbert Bátfai, Renátó Besenczi, Gergő Bogacsovics, Fanny Monori

In this article, we introduce a new conception of a family of esport games called Samu Entropy to try to improve dataflow program graphs like the ones that are based on Google's TensorFlow. Currently, the Samu Entropy pr…

Cross-Entropy Games and Frost Training

2026-05-26 · Arthur Renard, Franck Gabriel, Valentin Hartmann, Clément Hongler arxiv

We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy Games. The key idea is to exploit the gradient of the reward functio…