paper-with-me

홈 › Papers

A Long Way to Go: Investigating Length Correlations in RLHF

2023-10-05 · Prasann Singhal, Tanya Goyal, Jiacheng Xu, Greg Durrett

Great success has been reported using Reinforcement Learning from Human Feedback (RLHF) to align large language models, with open preference datasets enabling wider experimentation, particularly for "helpfulness" in tasks like dialogue and web question answering. Alongside these improvements, however, RLHF also often drives models to produce longer outputs. This paper demonstrates, on three diverse settings, that optimizing for response length is, much more than previously thought, a significant factor behind RLHF. Studying the strategies RL optimization uses to maximize reward, we find improvements in reward to largely be driven by increasing response length, instead of other features. Indeed, we find that even a purely length-based reward reproduces most downstream RLHF improvements over supervised fine-tuned models. Testing a comprehensive set of length-countering interventions, we identify the dominant source of these biases to be reward models, which, by studying training dynamics, we find are non-robust and easily influenced by length biases in preference data.

📄 PDF Abstract BibTeX arXiv:2310.03716

Code (1)

prasanns/rlhf-length-biases 공식 구현 pytorch

Tasks

Question Answering

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Mitigating Length Bias in RLHF through a Causal Lens

2025-11-16 · Hyeonji Kim, Sujeong Oh, Sanghack Lee arxiv

Reinforcement learning from human feedback (RLHF) is widely used to align large language models (LLMs) with human preferences. However, RLHF-trained reward models often exhibit length bias -- a systematic tendency to fav…

Reinforcement LearningData Augmentation

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

2025-01-16 · Chaoqi Wang, Zhuokai Zhao, Yibo Jiang, Zhaorun Chen 외

Recent advances in large language models (LLMs) have demonstrated significant progress in performing complex tasks. While Reinforcement Learning from Human Feedback (RLHF) has been effective in aligning LLMs with human p…

Causal InferencecounterfactualFairnessLanguage Modeling+2

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

2025-05-19 · Kangwen Zhao, JianFeng Cai, Jinhua Zhu, Ruopei Sun 외

Reinforcement Learning from Human Feedback relies on reward models to align large language models with human preferences. However, RLHF often suffers from reward hacking, wherein policy learning exploits flaws in the tra…

Relation

Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance

2025-12-29 · Zhuo Li, Pengyu Cheng, Zhechao Yu, Feifei Tong 외 arxiv

Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align large language models (LLMs) with human values. However, RM training data is commonly recognized as low-quality, containing …

Reinforcement Learning

OPPO: Accelerating PPO-based RLHF via Pipeline Overlap

2025-09-30 · Kaizhuo Yan, Yingjie Yu, Yifan Yu, Haizhong Zheng 외 arxiv

Proximal Policy Optimization (PPO)-based reinforcement learning from human feedback (RLHF) is a widely adopted paradigm for aligning large language models (LLMs) with human preferences. However, its training pipeline suf…

Reinforcement Learning