paper-with-me

홈 › Papers

Imperfect World Models are Exploitable

2026-05-15 · Logan Mondal Bhamidipaty, Esmeralda S. Whitammer, David Abel, Mykel J. Kochenderfer, Subramanian Ramamoorthy arxiv

We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred over another while the environment's true transition model implies the reverse. We analogize our definition with a prior characterization of reward hacking but show that the associated proof of inevitability does not transfer to exploitation. To overcome this obstruction, we develop a general theory of reward hacking and model exploitation that proves that exploitation is essentially unavoidable on large policy sets and yields the corresponding claim for hacking as a special case. Unfortunately, we also find that the conditions that guarantee unhackability in finite policy sets have no counterpart that precludes exploitation. Consequently, we introduce a relaxed notion of exploitation and derive a safe horizon within which it can be avoided. Taken together, our results establish a formal bridge between reward hacking and model exploitation and elucidate the limits of safe planning in world models.

📄 PDF Abstract BibTeX arXiv:2605.15960

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Playing the Player: A Heuristic Framework for Adaptive Poker AI

2025-12-04 · Andrew Paterson, Carl Sanders arxiv

For years, the discourse around poker AI has been dominated by the concept of solvers and the pursuit of unexploitable, machine-perfect play. This paper challenges that orthodoxy. It presents Patrick, an AI built on the …

Discovering Imperfectly Observable Adversarial Actions using Anomaly Detection

2020-04-22 · Olga Petrova, Karel Durkota, Galina Alperovich, Karel Horak 외

Anomaly detection is a method for discovering unusual and suspicious behavior. In many real-world scenarios, the examined events can be directly linked to the actions of an adversary, such as attacks on computer networks…

Anomaly Detection

Learning to Generate Cross-Task Unexploitable Examples

2025-12-15 · Haoxuan Qu, Qiuchi Xiang, Yujun Cai, Yirui Wu 외 arxiv

Unexploitable example generation aims to transform personal images into their unexploitable (unlearnable) versions before they are uploaded online, thereby preventing unauthorized exploitation of online personal images. …

AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play

2026-05-09 · Vlad Murgoci, Matthijs Spaan, Yaniv Oren arxiv

Poker is an imperfect information game that has served as a long-standing benchmark for decision-making under uncertainty. To maximize utility beyond the Nash equilibrium, an agent can deviate from Nash-equilibrium polic…

Mixture of Public and Private Distributions in Imperfect Information Games

2024-05-23 · Jérôme Arjonilla, Abdallah Saffidine, Tristan Cazenave

In imperfect information games (e.g. Bridge, Skat, Poker), one of the fundamental considerations is to infer the missing information while at the same time avoiding the disclosure of private information. Disregarding the…