paper-with-me

홈 › Papers

Auditing Reward Hackability in Code RL Training Environments

2026-06-14 · Shreshth Rajan arxiv

We measure the rate at which code RL environments accept incorrect solutions as correct. On a 49-task sample of SWE-bench Verified, 28.5% of tasks have test suites weak enough that a Docker-verified incorrect patch passes them. On 20 R2E-Gym tasks across 6 repositories, the same pipeline at single-shot exploit generation yields 25.0%. A random-effects meta-analysis over 134 frontier model submissions to SWE-bench Verified finds, within the same human-rated difficulty stratum, model Pass@1 is +14.14 percentage points higher on flagged-hackable tasks than on robust ones (95% CI [+11.80, +16.48]; one-sided p < 10^-6; I^2 = 0%; 123 of 134 models positive). We then describe a procedure for hardening the broken tasks. An inline LLM judge with a Docker gold-sanity gate runs each generated test against the gold solution before the judge is consulted. On the 11 broken tasks in the audit, the gate flags 65 of 105 decisive LLM-generated tests as failing on the gold patch itself, a 61.9% per-augmentation defect rate the LLM judge alone misses. With diversity-biased retry, the loop converges 9 of 11 tasks to a gated upgrade.

📄 PDF Abstract BibTeX arXiv:2606.16062

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Defining and Characterizing Reward Hacking

2022-09-27 · Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David Krueger

We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function, $\mathcal{\tilde{R}}$, leads to poor performance according to the true reward function, $\mathca…

Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models

2026-02-20 · Rishabh Tiwari, Aditya Tomar, Udbhav Bamba, Monishwaran Maheswaran 외 arxiv

Process Reward Models (PRMs) are rapidly becoming the backbone of LLM reasoning pipelines, yet we demonstrate that state-of-the-art PRMs are systematically exploitable under adversarial optimization pressure. To address …

Alignment Risks from Capability-Seeking RL Training

2026-02-12 · Yujun Zhou, Yue Huang, Han Bao, Kehan Guo 외 arxiv

While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments. We investigate whether l…

Reinforcement Learning

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards

2025-11-18 · Yule Liu, Heyi Zhang, Jinyi Zheng, Zhen Sun 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on non-public, high-value prompt sets raises concerns about unauthorized data us…

Reinforcement Learning

Imperfect World Models are Exploitable

2026-05-15 · Logan Mondal Bhamidipaty, Esmeralda S. Whitammer, David Abel, Mykel J. Kochenderfer 외 arxiv

We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred over another while the environment's true…

Reinforcement Learning