paper-with-me

홈 › Papers

One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models

2026-02-06 · Daniel Fein, Max Lamparth, Violet Xiang, Mykel J. Kochenderfer, Nick Haber arxiv

Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, including the state-of-the-art, we find that issues persist despite prior work with respect to length, sycophancy, and overconfidence. We also discover new issues related to bias toward model-specific ``styles'' and answer-order. We categorize RM failures as tractable or resistant to linear intervention and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. Our proposed mechanistic reward shaping reduces targeted biases without degrading reward quality and while using minimal labeled data. The method is extensible to new biases, model-internal, and generalizes out-of-distribution.

📄 PDF Abstract BibTeX arXiv:2603.03291

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge

2024-10-02 · Xiefeng Wu

Q-shaping is an extension of Q-value initialization and serves as an alternative to reward shaping for incorporating domain knowledge to accelerate agent training, thereby improving sample efficiency by directly shaping …

Language ModelingLanguage ModellingLarge Language Model

Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping

2020-11-05 · NeurIPS 2020 12 · Yujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 외

Reward shaping is an effective technique for incorporating domain knowledge into reinforcement learning (RL). Existing approaches such as potential-based reward shaping normally make full use of a given shaping reward fu…

MuJoCoReinforcement Learning (RL)

On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

2026-07-03 · Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi 외 arxiv

Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures. At test time, however, the goal is not to sample from this prior, but to use a limited o…

Unbiased learning with State-Conditioned Rewards in Adversarial Imitation Learning

2021-01-01 · Dong-Sig Han, Hyunseo Kim, Hyundo Lee, Je-Hwan Ryu 외

Adversarial imitation learning has emerged as a general and scalable framework for automatic reward acquisition. However, we point out that previous methods commonly exploited occupancy-dependent reward learning formulat…

continuous-controlContinuous ControlImitation Learningreinforcement-learning+2

Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping

2025-03-26 · Mohammad Saif Nazir, Chayan Banerjee

Reinforcement learning often faces challenges with reward misalignment, where agents optimize for given rewards but fail to exhibit the desired behaviors. This occurs when the reward function incentivizes proxy behaviors…

continuous-controlContinuous Control