paper-with-me

홈 › Papers

Understanding the Effects of RLHF on LLM Generalisation and Diversity

2023-10-10 · Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, Roberta Raileanu

Large language models (LLMs) fine-tuned with reinforcement learning from human feedback (RLHF) have been used in some of the most widely deployed AI models to date, such as OpenAI's ChatGPT or Anthropic's Claude. While there has been significant work developing these methods, our understanding of the benefits and downsides of each stage in RLHF is still limited. To fill this gap, we present an extensive analysis of how each stage of the process (i.e. supervised fine-tuning (SFT), reward modelling, and RLHF) affects two key properties: out-of-distribution (OOD) generalisation and output diversity. OOD generalisation is crucial given the wide range of real-world scenarios in which these models are being used, while output diversity refers to the model's ability to generate varied outputs and is important for a variety of use cases. We perform our analysis across two base models on both summarisation and instruction following tasks, the latter being highly relevant for current LLM use cases. We find that RLHF generalises better than SFT to new inputs, particularly as the distribution shift between train and test becomes larger. However, RLHF significantly reduces output diversity compared to SFT across a variety of measures, implying a tradeoff in current LLM fine-tuning methods between generalisation and diversity. Our results provide guidance on which fine-tuning method should be used depending on the application, and show that more research is needed to improve the tradeoff between generalisation and diversity.

📄 PDF Abstract BibTeX arXiv:2310.06452

Code (1)

facebookresearch/rlfh-gen-div 공식 구현 pytorch

Tasks

DiversityInstruction Following

Methods 이 논문이 사용한 방법론

SFT Shrink and Fine-Tune, or SFT, is a type of distillation that avoids explicit distillation by copying parameters to a student student model and then fine-tuning.…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Generalisation of RLHF under Reward Shift and Clipped KL Regularisation

2026-02-25 · Kenton Tang, Yuzhu Chen, Fengxiang He arxiv

Alignment and adaptation in large language models heavily rely on reinforcement learning from human feedback (RLHF); yet, theoretical understanding of its generalisability remains premature, especially when the learned r…

Reinforcement Learning

Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

2025-03-28 · Wei Shen, Guanlin Liu, Zheng Wu, Ruofei Zhu 외

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models with human preferences. While recent research has focused on algorithmic improvements, the importance of prompt-data constru…

Diversity

Curiosity-Driven Reinforcement Learning from Human Feedback

2025-01-20 · Haoran Sun, Yekun Chai, Shuohuan Wang, Yu Sun 외

Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but often at the cost of reduced output diversity. This trade-off between diversity …

DiversityInstruction Followingreinforcement-learningReinforcement Learning+1

Understanding Likelihood Over-optimisation in Direct Alignment Algorithms

2024-10-15 · Zhengyan Shi, Sander Land, Acyr Locatelli, Matthieu Geist 외

Direct Alignment Algorithms (DAAs), such as Direct Preference Optimisation (DPO) and Identity Preference Optimisation (IPO), have emerged as alternatives to online Reinforcement Learning from Human Feedback (RLHF) algori…

Diversity

Perspectives on the Social Impacts of Reinforcement Learning with Human Feedback

2023-03-06 · Gabrielle Kaili-May Liu

Is it possible for machines to think like humans? And if it is, how should we go about teaching them to do so? As early as 1950, Alan Turing stated that we ought to teach machines in the way of teaching a child. Reinforc…

Misinformationreinforcement-learningReinforcement Learning (RL)