paper-with-me

홈 › Papers

Preference Optimization by Estimating the Ratio of the Data Distribution

2025-05-26 · Yeongmin Kim, HeeSun Bae, Byeonghu Na, Il-Chul Moon

Direct preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a generalized DPO loss that enables a policy model to match the target policy from a likelihood ratio estimation perspective. The ratio of the target policy provides a unique identification of the policy distribution without relying on reward models or partition functions. This allows the generalized loss to retain both simplicity and theoretical guarantees, which prior work such as $f$-PO fails to achieve simultaneously. We propose Bregman preference optimization (BPO), a generalized framework for ratio matching that provides a family of objective functions achieving target policy optimality. BPO subsumes DPO as a special case and offers tractable forms for all instances, allowing implementation with a few lines of code. We further develop scaled Basu's power divergence (SBA), a gradient scaling method that can be used for BPO instances. The BPO framework complements other DPO variants and is applicable to target policies defined by these variants. In experiments, unlike other probabilistic loss extensions such as $f$-DPO or $f$-PO, which exhibit a trade-off between generation fidelity and diversity, instances of BPO improve both win rate and entropy compared with DPO. When applied to Llama-3-Instruct-8B, BPO achieves state-of-the-art performance among Llama-3-8B backbones, with a 55.9\% length-controlled win rate on AlpacaEval2.

📄 PDF Abstract BibTeX arXiv:2505.19601

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

DPO 설명 없음

Similar Papers 제목 키워드 기반

Preference Models assume Proportional Hazards of Utilities

2025-08-15 · Chirag Nagpal arxiv

Approaches for estimating preferences from human annotated data typically involves inducing a distribution over a ranked list of choices such as the Plackett-Luce model. Indeed, modern AI alignment tools such as Reward M…

Accelerated Preference Optimization for Large Language Model Alignment

2024-10-08 · Jiafan He, Huizhuo Yuan, Quanquan Gu

Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal tool for aligning large language models (LLMs) with human preferences. Direct Preference Optimization (DPO), one of the most popular approaches, …

Language ModelingLanguage ModellingLarge Language Model

No Preference Left Behind: Group Distributional Preference Optimization

2024-12-28 · Binwei Yao, Zefan Cai, Yun-Shiuan Chuang, Shanglin Yang 외

Preferences within a group of people are not uniform but follow a distribution. While existing alignment methods like Direct Preference Optimization (DPO) attempt to steer models to reflect human preferences, they strugg…

DiversityLanguage ModelingLanguage Modelling

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

2025-06-03 · Yunhong Lu, Qichao Wang, Hengyuan Cao, Xiaoyin Xu 외

Direct Preference Optimization (DPO) aligns text-to-image (T2I) generation models with human preferences using pairwise preference data. Although substantial resources are expended in collecting and labeling datasets, a …

Talos: Optimizing Top-$K$ Accuracy in Recommender Systems

2026-01-27 · Shengjia Zhang, Weiqin Yang, Jiawei Chen, Peng Wu 외 arxiv

Recommender systems (RS) aim to retrieve a small set of items that best match individual user preferences. Naturally, RS place primary emphasis on the quality of the Top-$K$ results rather than performance across the ent…