paper-with-me

홈 › Papers

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking

2026-04-29 · Disha Singha arxiv

Reinforcement learning from human feedback (RLHF) systems face a compounding alignment challenge: not only are learned reward models uncertain about unseen state-action pairs, but the human preference annotations they are trained on are themselves inconsistent, context-dependent, and noisy. Existing approaches address these uncertainty sources in isolation - epistemic uncertainty is used to guide exploration, while preference uncertainty is absorbed during reward model training but discarded during policy optimization. We introduce Uncertainty-Aware Reward Discounting (UARD), a principled framework that jointly models epistemic uncertainty in value estimation via ensemble disagreement and aleatoric uncertainty in human preference annotations via annotator variability, combining these signals through a confidence-adjusted Reliability Filter that adaptively modulates reward weighting during policy optimization. We prove that this dynamic discounting preserves the contraction property of the Bellman operator, guaranteeing convergence to a unique fixed point, and provide an information-theoretic justification grounded in the Information Bottleneck principle. Empirically, UARD reduces reward hacking incidents by up to 93.6% across discrete decision-making and continuous control benchmarks (MuJoCo) compared to nine baselines including DQN, Ensemble-DQN, CQL, CPO, TRPO, SAC, EDAC, SUNRISE, and PPO, while maintaining competitive task performance on well-specified rewards. Under annotation noise ranging from 10% to 30% Gaussian perturbation, UARD retains near-zero safety violations compared to baselines' near-linear degradation. These results demonstrate that treating uncertainty as an active component of the optimization objective - rather than a passive diagnostic signal - provides a principled pathway toward more reliable and aligned RL systems.

📄 PDF Abstract BibTeX arXiv:2604.26360

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningContinuous Control

Similar Papers 제목 키워드 기반

\$1 Today or \$2 Tomorrow? The Answer is in Your Facebook Likes

2017-03-22 · Tao Ding, Warren K. Bickel, SHimei Pan

In economics and psychology, delay discounting is often used to characterize how individuals choose between a smaller immediate reward and a larger delayed reward. People with higher delay discounting rate (DDR) often ch…

Decision Making

Awareness of self-control

2024-02-16 · Mohammad Mehdi Mousavi, Mahdi Kohan Sefidi, Shirin Allahyarkhani

Economists modeled self-control problems in decisions of people with the time-inconsistence preferences model. They argued that the source of self-control problems could be uncertainty and temptation. This paper uses an …

Marketing

UCPO: Uncertainty-Aware Policy Optimization

2026-01-30 · Xianzhou Zeng, Jing Huang, Chunmei Xie, Gongrui Nan 외 arxiv

The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitigating overconfident errors in high-stakes applications. However, existing…

Mathematical Reasoning

Decision-making and Fuzzy Temporal Logic

2019-01-07 · José Cláudio do Nascimento

This paper shows that the fuzzy temporal logic can model figures of thought to describe decision-making behaviors. In order to exemplify, some economic behaviors observed experimentally were modeled from problems of choi…

Decision Making

Discounting and Drug Seeking in Biological Hierarchical Reinforcement Learning

2025-06-05 · Vardhan Palod, Pranav Mahajan, Veeky Baths, Boris S. Gutkin

Despite a strong desire to quit, individuals with long-term substance use disorder (SUD) often struggle to resist drug use, even when aware of its harmful consequences. This disconnect between knowledge and compulsive be…

Hierarchical Reinforcement Learningreinforcement-learningReinforcement Learning