paper-with-me

Papers

Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

2026-03-07 · Muhammad Khalifa, Zohaib Khan, Omer Tafveez, Hao Peng, Lu Wang arxiv

Reward hacking is a form of misalignment in which models overoptimize proxy rewards without genuinely solving the underlying task. Precisely measuring reward hacking occurrence remains challenging because true task rewards are often expensive or impossible to compute. We introduce Countdown-Code, a minimal environment where models can both solve a mathematical reasoning task and manipulate the test harness. This dual-access design creates a clean separation between proxy rewards (test pass/fail) and true rewards (mathematical correctness), enabling accurate measurement of reward-hacking rates. Using this environment, we study reward hacking in open-weight LLMs and find that such behaviors can be unintentionally learned during supervised fine-tuning (SFT) when even a small fraction of reward-hacking trajectories leak into training data. As little as 1\% contamination in distillation SFT data is sufficient for models to internalize reward hacking which resurfaces during subsequent reinforcement learning (RL). We further show that RL amplifies misalignment and drives its generalization beyond the original domain. We open-source our environment and code to facilitate future research on reward hacking in LLMs. Our results reveal a previously underexplored pathway through which reward hacking can emerge and persist in LLMs, underscoring the need for more rigorous validation of synthetic SFT data. Code is available at https://github.com/zohaib-khan5040/Countdown-Code.

📄 PDF Abstract BibTeX arXiv:2603.07084

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

2026-01-28 · Minjae Kwon, Josephine Lamp, Lu Feng arxiv

Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safety guarantees transfer to deployment under distribution shift, using diabetes …

Reinforcement Learning

How Does RL Post-training Induce Skill Composition? A Case Study on Countdown

2025-12-01 · Simon Park, Simran Kaur, Sanjeev Arora arxiv

While reinforcement learning (RL) successfully enhances reasoning in large language models, its role in fostering compositional generalization (the ability to synthesize novel skills from known components) is often confl…

Reinforcement Learning

COUNTDOWN: Contextually Sparse Activation Filtering Out Unnecessary Weights in Down Projection

2025-05-23 · Jaewon Cheon, Pilsung Kang

The growing size of large language models has created significant computational inefficiencies. To address this challenge, sparse activation methods selectively deactivates non-essential parameters during inference, redu…

ALICE: An Interpretable Neural Architecture for Generalization in Substitution Ciphers

2025-09-08 · Jeff Shen, Lindsay M. Smith arxiv

We present cryptogram solving as an ideal testbed for studying neural network reasoning and generalization; models must decrypt text encoded with substitution ciphers, choosing from 26! possible mappings without explicit…

The (Final) countdown

2015-02-19 · Jean-Marc Alliot

The Countdown game is one of the oldest TV show running in the world. It started broadcasting in 1972 on the french television and in 1982 on British channel 4, and it has been running since in both countries. The game, …