paper-with-me

홈 › Papers

Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

2024-02-11 · Chenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong, Nan Jiang, Tong Zhang

We investigate Reinforcement Learning from Human Feedback (RLHF) in the context of a general preference oracle. In particular, we do not assume the existence of a reward function and an oracle preference signal drawn from the Bradley-Terry model as most of the prior works do. We consider a standard mathematical formulation, the reverse-KL regularized minimax game between two LLMs for RLHF under general preference oracle. The learning objective of this formulation is to find a policy so that it is consistently preferred by the KL-regularized preference oracle over any competing LLMs. We show that this framework is strictly more general than the reward-based one, and propose sample-efficient algorithms for both the offline learning from a pre-collected preference dataset and online learning where we can query the preference oracle along the way of training. Empirical studies verify the effectiveness of the proposed framework.

📄 PDF Abstract BibTeX arXiv:2402.07314

Code (1)

weixiongust/rlhf-reward-modeling 공식 구현 pytorch

Similar Papers 제목 키워드 기반

RLHF Workflow: From Reward Modeling to Online RLHF

2024-05-13 · Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang 외

We present the workflow of Online Iterative Reinforcement Learning from Human Feedback (RLHF) in this technical report, which is widely reported to outperform its offline counterpart by a large margin in the recent large…

ChatbotHumanEvalLanguage ModellingLarge Language Model+1

SAIL: Self-Improving Efficient Online Alignment of Large Language Models

2024-06-21 · Mucong Ding, Souradip Chakraborty, Vibhu Agrawal, Zora Che 외

Reinforcement Learning from Human Feedback (RLHF) is a key method for aligning large language models (LLMs) with human preferences. However, current offline alignment approaches like DPO, IPO, and SLiC rely heavily on fi…

Bilevel Optimization

PILAF: Optimal Human Preference Sampling for Reward Modeling

2025-02-06 · Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng, Julia Kempe 외

As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique, translating prefer…

Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

2024-06-30 · Yuheng Zhang, Dian Yu, Baolin Peng, Linfeng Song 외

Reinforcement Learning with Human Feedback (RLHF) has achieved great success in aligning large language models (LLMs) with human preferences. Prevalent RLHF approaches are reward-based, following the Bradley-Terry (BT) m…

Online Preference Alignment for Language Models via Count-based Exploration

2025-01-22 · Chenjia Bai, Yang Zhang, Shuang Qiu, Qiaosheng Zhang 외

Reinforcement Learning from Human Feedback (RLHF) has shown great potential in fine-tuning Large Language Models (LLMs) to align with human preferences. Existing methods perform preference alignment from a fixed dataset,…

Instruction Following