paper-with-me

Papers

Online Self-Preferring Language Models

2024-05-23 · Yuanzhao Zhai, Zhuo Zhang, Kele Xu, Hanyang Peng, Yue Yu, Dawei Feng, Cheng Yang, Bo Ding, Huaimin Wang

Aligning with human preference datasets has been critical to the success of large language models (LLMs). Reinforcement learning from human feedback (RLHF) employs a costly reward model to provide feedback for on-policy sampling responses. Recently, offline methods that directly fit responses with binary preferences in the dataset have emerged as alternatives. However, existing methods do not explicitly model preference strength information, which is crucial for distinguishing different response pairs. To overcome this limitation, we propose Online Self-Preferring (OSP) language models to learn from self-generated response pairs and self-judged preference strengths. For each prompt and corresponding self-generated responses, we introduce a ranked pairing method to construct multiple response pairs with preference strength information. We then propose the soft-preference cross-entropy loss to leverage such information. Empirically, we demonstrate that leveraging preference strength is crucial for avoiding overfitting and enhancing alignment performance. OSP achieves state-of-the-art alignment performance across various metrics in two widely used human preference datasets. OSP is parameter-efficient and more robust than the dominant online method, RLHF when limited offline data are available and generalizing to out-of-domain tasks. Moreover, OSP language models established by LLMs with proficiency in self-preferring can efficiently self-improve without external supervision.

📄 PDF Abstract BibTeX arXiv:2405.14103

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

2026-01-30 · Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas 외 arxiv

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of automated post-training and evaluation workflows. However, it is difficult to disent…

Differentially Private Learning Needs Better Model Initialization and Self-Distillation

2024-10-23 · Ivoline C. Ngong, Joseph P. Near, Niloofar Mireshghallah

Differentially private SGD (DPSGD) enables privacy-preserving training of language models, but often reduces utility, diversity, and linguistic quality. We introduce DPRefine, a three-phase method that initializes a mode…

DiversityPrivacy Preserving

Mitigating Self-Preference by Authorship Obfuscation

2025-12-05 · Taslim Mahbub, Shi Feng arxiv

Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in evaluations. One such bias is self-prefe…

Early Gains Matter: A Case for Preferring Generative over Discriminative Crowdsourcing Models

2015-05-01 · HLT 2015 5 · Eric Ringger, Paul Felt, Kevin Seppi, Robbie Haertel 외
text-classificationText ClassificationVariational Inference

The Role of Verb Semantics in Hungarian Verb-Object Order

2020-06-16 · Dorottya Demszky, László Kálmán, Dan Jurafsky, Beth Levin

Hungarian is often referred to as a discourse-configurational language, since the structural position of constituents is determined by their logical function (topic or comment) rather than their grammatical function (e.g…

Object