paper-with-me

홈 › Papers

Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

2024-10-01 · Xingzhou Lou, Dong Yan, Wei Shen, Yuzi Yan, Jian Xie, Junge Zhang

Reward models (RM) play a critical role in aligning generations of large language models (LLM) to human expectations. However, prevailing RMs fail to capture the stochasticity within human preferences and cannot effectively evaluate the reliability of reward predictions. To address these issues, we propose Uncertain-aware RM (URM) and Uncertain-aware RM Ensemble (URME) to incorporate and manage uncertainty in reward modeling. URM can model the distribution of disentangled attributes within human preferences, while URME quantifies uncertainty through discrepancies in the ensemble, thereby identifying potential lack of knowledge during reward evaluation. Experiment results indicate that the proposed URM achieves state-of-the-art performance compared to models with the same size, demonstrating the effectiveness of modeling uncertainty within human preferences. Furthermore, empirical results show that through uncertainty quantification, URM and URME can identify unreliable predictions to improve the quality of reward evaluations.

📄 PDF Abstract BibTeX arXiv:2410.00847

Code (0)

등록된 구현이 없습니다.

Tasks

Uncertainty Quantification

Similar Papers 제목 키워드 기반

UCPO: Uncertainty-Aware Policy Optimization

2026-01-30 · Xianzhou Zeng, Jing Huang, Chunmei Xie, Gongrui Nan 외 arxiv

The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitigating overconfident errors in high-stakes applications. However, existing…

Mathematical Reasoning

SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards

2026-02-24 · Dengjia Zhang, Xiaoou Liu, Lu Cheng, Yaqing Wang 외 arxiv

Large language models (LLMs) are increasingly deployed as multi-step decision-making agents, where effective reward design is essential for guiding learning. Although recent work explores various forms of reward shaping …

Reinforcement Learning

RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models

2026-02-27 · Daniel Yang, Samuel Stante, Florian Redhardt, Lena Libon 외 arxiv

Reward models are central to aligning large language models (LLMs) with human preferences. Yet most approaches rely on pointwise reward estimates that overlook the epistemic uncertainty in reward models arising from limi…

Active Learning

Uncertainty-Aware Reward Modeling for Stable RLHF

2026-06-18 · Licheng Pan, Haocheng Yang, Haoxuan Li, Yichen Sun 외 arxiv

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. However, this pipeline faces two fundamen…

Reinforcement Learning

Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling

2024-10-15 · Guiyu Zhang, Huan-ang Gao, Zijian Jiang, Hao Zhao 외

In this paper, we focus on the task of conditional image generation, where an image is synthesized according to user instructions. The critical challenge underpinning this task is ensuring both the fidelity of the genera…

Conditional Image GenerationImage Generation