paper-with-me

홈 › Papers

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment

2026-05-13 · Ye Wang, Jing Liu, Toshiaki Koike-Akino arxiv

Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptation as alignment objectives and reward targets evolve. Existing theoretical analyses justify these methods as approximations to sampling from distributions optimally tilted toward a given reward model. We extend these techniques by introducing reference-model temperature adjustment, which leads to further generalization of inference-time alignment to ensembles of generative reward models combined as a sharpened logarithmic opinion pool (SLOP). To mitigate reward hacking, we propose an algorithm for calibrating SLOP weight parameters and experimentally demonstrate that it improves robustness while preserving alignment performance.

📄 PDF Abstract BibTeX arXiv:2605.13537

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Are we really tilting? The mechanics of reward guidance in flow and diffusion models

2026-06-01 · Sanjit Dandapanthula, Nicholas M. Boffi arxiv

Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these methods are prone to reward hacking: the guided model over-optimizes the…

Text-to-Image Generation

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Trainig-Time Reward Hacking in Code Generation

2026-04-26 · Lichen Li, Hengguang Zhou, Yijun Liang, Tianyi Zhou 외 arxiv

Reward hacking in code generation, where models exploit evaluation loopholes to obtain high reward without correctly solving the intended task, poses a critical challenge for Reinforcement Learning (RL) and the deploymen…

Reinforcement LearningCode Generation

How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis

2026-05-23 · Rei Higuchi, Ryotaro Kawata, Akifumi Wachi, Shokichi Takakura 외 arxiv

Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depends on errors in reward-tilted regions. …

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

2026-06-03 · Xuekang Wang, Zhuoyuan Hao, Shuo Hou, Hao Peng 외 arxiv

Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and …

Reinforcement Learning

Spontaneous Reward Hacking in Iterative Self-Refinement

2024-07-05 · Jane Pan, He He, Samuel R. Bowman, Shi Feng

Language models are capable of iteratively improving their outputs based on natural language feedback, thus enabling in-context optimization of user preference. In place of human users, a second language model can be use…

Language ModelingLanguage Modelling