paper-with-me

홈 › Papers

Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison

2024-09-15 · Judy Hanwen Shen, Archit Sharma, Jun Qin

The goal of aligning language models to human preferences requires data that reveal these preferences. Ideally, time and money can be spent carefully collecting and tailoring bespoke preference data to each downstream application. However, in practice, a select few publicly available preference datasets are often used to train reward models for reinforcement learning from human feedback (RLHF). While new preference datasets are being introduced with increasing frequency, there are currently no existing efforts to measure and compare these datasets. In this paper, we systematically study preference datasets through three perspectives: scale, label noise, and information content. We propose specific metrics for each of these perspectives and uncover different axes of comparison for a better understanding of preference datasets. Our work is a first step towards a data-centric approach to alignment by providing perspectives that aid in training efficiency and iterative data collection for RLHF.

📄 PDF Abstract BibTeX arXiv:2409.09603

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How to Evaluate Reward Models for RLHF

2024-10-18 · Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang 외

We introduce a new benchmark for reward models that quantifies their ability to produce strong language models through RLHF (Reinforcement Learning from Human Feedback). The gold-standard approach is to run a full RLHF t…

Reward Difference Optimization For Sample Reweighting In Offline RLHF

2024-08-18 · Shiqi Wang, Zhengze Zhang, Rui Zhao, Fei Tan 외

With the rapid advances in Large Language Models (LLMs), aligning LLMs with human preferences become increasingly important. Although Reinforcement Learning with Human Feedback (RLHF) proves effective, it is complicated …

PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

2024-06-20 · Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 외

In this work, we introduce the PKU-SafeRLHF dataset, designed to promote research on safety alignment in large language models (LLMs). As a sibling project to SafeRLHF and BeaverTails, we separate annotations of helpfuln…

Question AnsweringSafety Alignment

Continual SFT Matches Multimodal RLHF with Negative Supervision

2024-11-22 · CVPR 2025 1 · Ke Zhu, Yu Wang, Yanpeng Sun, Qiang Chen 외

Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiority over continual SFT during this prefer…

FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback

2023-07-20 · Ashish Singh, Prateek Agarwal, Zixuan Huang, Arpita Singh 외

Captions are crucial for understanding scientific visualizations and documents. Existing captioning methods for scientific figures rely on figure-caption pairs extracted from documents for training, many of which fall sh…

Caption Generation