paper-with-me

홈 › Papers

A density estimation perspective on learning from pairwise human preferences

2023-11-23 · Vincent Dumoulin, Daniel D. Johnson, Pablo Samuel Castro, Hugo Larochelle, Yann Dauphin

Learning from human feedback (LHF) -- and in particular learning from pairwise preferences -- has recently become a crucial ingredient in training large language models (LLMs), and has been the subject of much research. Most recent works frame it as a reinforcement learning problem, where a reward function is learned from pairwise preference data and the LLM is treated as a policy which is adapted to maximize the rewards, often under additional regularization constraints. We propose an alternative interpretation which centers on the generative process for pairwise preferences and treats LHF as a density estimation problem. We provide theoretical and empirical results showing that for a family of generative processes defined via preference behavior distribution equations, training a reward function on pairwise preferences effectively models an annotator's implicit preference distribution. Finally, we discuss and present findings on "annotator misspecification" -- failure cases where wrong modeling assumptions are made about annotator behavior, resulting in poorly-adapted models -- suggesting that approaches that learn from pairwise human preferences could have trouble learning from a population of annotators with diverse viewpoints.

📄 PDF Abstract BibTeX arXiv:2311.14115

Code (1)

google-deepmind/pbde 공식 구현

Tasks

Density Estimation

Similar Papers 제목 키워드 기반

Score-Based Density Estimation from Pairwise Comparisons

2025-10-10 · Petrus Mikkola, Luigi Acerbi, Arto Klami arxiv

We study density estimation from pairwise comparisons, motivated by expert knowledge elicitation and learning from human feedback. We relate the unobserved target density to a tempered winner density (marginal density of…

Density Estimation

Prediction-Powered Ranking of Large Language Models

2024-02-27 · Ivi Chatzi, Eleni Straitouri, Suhas Thejaswi, Manuel Gomez Rodriguez

Large language models are often ranked according to their level of alignment with human preferences -- a model is better than other models if its outputs are more frequently preferred by humans. One of the popular ways t…

ChatbotLanguage ModellingLarge Language ModelPrediction

Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment

2026-04-06 · Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai, Sekitoshi Kanai 외 arxiv

Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models such as the Bradley-Terry model, this assu…

Estimating Summary Quality with Pairwise Preferences

2018-06-01 · NAACL 2018 6 · Markus Zopf

Automatic evaluation systems in the field of automatic summarization have been relying on the availability of gold standard summaries for over ten years. Gold standard summaries are expensive to obtain and often require …

Text Summarization

The quality of priority ratios estimation in relation to a selected prioritization procedure and consistency measure for a Pairwise Comparison Matrix

2017-04-06 · Paul Thaddeus Kazibudzki

An overview of current debates and contemporary research devoted to the modeling of decision making processes and their facilitation directs attention to the Analytic Hierarchy Process (AHP). At the core of the AHP are v…

Decision Making