paper-with-me

Papers

Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?

2025-08-14 · Zetian Sun, Dongfang Li, Xuhui Chen, Baotian Hu, Min Zhang arxiv

The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize the expected reward that reflects human preferences. Recently, Direct Preference Optimization~(DPO) was proposed as a LM alignment method that directly optimize the policy from static preference data, and further improved by incorporating on-policy sampling~(i.e., preference candidates generated during the training loop) for better LM alignment. However, we show on-policy data is not always optimal, with systematic effectiveness difference emerging between static and on-policy preference candidates. For example, on-policy data can result in a $3\times$ effectiveness compared with static data for Llama-3, and a $0.4\times$ effectiveness for Zephyr. To explain the phenomenon, we propose the alignment stage assumption, which divides the alignment process into two distinct stages: the preference injection stage, which benefits from diverse data, and the preference fine-tuning stage, which favors high-quality data. Through theoretical and empirical analysis, we characterize these stages and propose an effective algorithm to identify the boundaries between them. We perform experiments on $5$ models~(Llama, Zephyr, Phi-2, Qwen, Pythia) and $2$ alignment methods~(DPO, SLiC-HF) to show the generalizability of alignment stage assumption and the effectiveness of the boundary measurement algorithm.

📄 PDF Abstract BibTeX arXiv:2508.10530

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Best-Effort Policies for Robust Markov Decision Processes

2025-08-11 · Alessandro Abate, Thom Badings, Giuseppe De Giacomo, Francesco Fabiano arxiv

We study the common generalization of Markov decision processes (MDPs) with sets of transition probabilities, known as robust MDPs (RMDPs). A standard goal in RMDPs is to compute a policy that maximizes the expected retu…

Rediscovery

2025-04-28 · Martino Banchio, Suraj Malladi

We model search in settings where decision makers know what can be found but not where to find it. A searcher faces a set of choices arranged by an observable attribute. Each period, she either selects a choice and pays …

Attribute

Group Learning and Opinion Diffusion in a Broadcast Network

2013-09-14 · Yang Liu, Mingyan Liu

We analyze the following group learning problem in the context of opinion diffusion: Consider a network with $M$ users, each facing $N$ options. In a discrete time setting, at each time step, each user chooses $K$ out of…

Multi-Product Dynamic Pricing in High-Dimensions with Heterogeneous Price Sensitivity

2019-01-04 · Adel Javanmard, Hamid Nazerzadeh, Simeng Shao

We consider the problem of multi-product dynamic pricing, in a contextual setting, for a seller of differentiated products. In this environment, the customers arrive over time and products are described by high-dimension…

SensitivityVocal Bursts Intensity Prediction

Dynamic Pricing in High-dimensions

2016-09-24 · Adel Javanmard, Hamid Nazerzadeh

We study the pricing problem faced by a firm that sells a large number of products, described via a wide range of features, to customers that arrive over time. Customers independently make purchasing decisions according …

Vocal Bursts Intensity Prediction