paper-with-me

홈 › Papers

Why DPO is a Misspecified Estimator and How to Fix It

2025-10-23 · Aditya Gopalan, Sayak Ray Chowdhury, Debangshu Banerjee arxiv

Direct alignment algorithms such as Direct Preference Optimization (DPO) fine-tune models based on preference data, using only supervised learning instead of two-stage reinforcement learning with human feedback (RLHF). We show that DPO encodes a statistical estimation problem over reward functions induced by a parametric policy class. When the true reward function that generates preferences cannot be realized via the policy class, DPO becomes misspecified, resulting in failure modes such as preference order reversal, worsening of policy reward, and high sensitivity to the input preference data distribution. On the other hand, we study the local behavior of two-stage RLHF for a parametric class and relate it to a natural gradient step in policy space. Our fine-grained geometric characterization allows us to propose AuxDPO, which introduces additional auxiliary variables in the DPO loss function to help move towards the RLHF solution in a principled manner and mitigate the misspecification in DPO. We empirically demonstrate the superior performance of AuxDPO on didactic bandit settings as well as LLM alignment tasks.

📄 PDF Abstract BibTeX arXiv:2510.20413

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Generalized Automatic Least Squares: Efficiency Gains from Misspecified Heteroscedasticity Models

2023-04-14 · Bulat Gafarov

It is well known that in the presence of heteroscedasticity ordinary least squares estimator is not efficient. I propose a generalized automatic least squares estimator (GALS) that makes partial correction of heterosceda…

Non-Bayesian Post-Model-Selection Estimation as Estimation Under Model Misspecification

2023-08-22 · Nadav Harel, Tirza Routtenberg

In many parameter estimation problems, the exact model is unknown and is assumed to belong to a set of candidate models. In such cases, a predetermined data-based selection rule selects a parametric model from a set of c…

channel selectionmodelModel Selectionparameter estimation

Sequential prediction under log-loss and misspecification

2021-01-29 · Meir Feder, Yury Polyanskiy

We consider the question of sequential prediction under the log-loss in terms of cumulative regret. Namely, given a hypothesis class of distributions, learner sequentially predicts the (distribution of the) next letter i…

Density EstimationModel SelectionPrediction

Design-Robust Two-Way-Fixed-Effects Regression For Panel Data

2021-07-29 · Dmitry Arkhangelsky, Guido W. Imbens, Lihua Lei, Xiaoman Luo

We propose a new estimator for average causal effects of a binary treatment with panel data in settings with general treatment patterns. Our approach augments the popular two-way-fixed-effects specification with unit-spe…

regressionVocal Bursts Valence Prediction

Improving the Accuracy of Amortized Model Comparison with Self-Consistency

2025-08-28 · Šimon Kucharský, Aayush Mishra, Daniel Habermann, Stefan T. Radev 외 arxiv

Amortized Bayesian model comparison (BMC) enables fast probabilistic ranking of models via simulation-based training of neural surrogates. However, the accuracy of neural surrogates deteriorates when simulation models ar…