paper-with-me

Papers

An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

2026-01-09 · Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras arxiv

Preference tuning aligns base language models to human judgments of quality, helpfulness, or safety by optimizing over explicit preference signals rather than likelihood alone. Prior work has shown that preference tuning degrades performance and reduces helpfulness outside the training domain. However, the extent to which adaptation strategies mitigate this domain shift remains unexplored. We address this challenge by conducting a comprehensive and systematic study of alignment generalization under domain shift. We compare five popular alignment objectives and various adaptation strategies from source to target, including target-domain supervised fine-tuning and pseudo-labeling, across summarization, question-answering helpfulness, and safety alignment tasks. Our findings reveal systematic differences in generalization across alignment objectives under domain shift. We show that adaptation strategies based on pseudo-labeling substantially reduce domain-shift degradation but induce mode collapse, revealing a generalization-diversity trade-off.

📄 PDF Abstract BibTeX arXiv:2601.05882

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MallowsPO: Fine-Tune Your LLM with Preference Dispersions

2024-05-23 · Haoxian Chen, Hanyang Zhao, Henry Lam, David Yao 외

Direct Preference Optimization (DPO) has recently emerged as a popular approach to improve reinforcement learning with human feedback (RLHF), leading to better techniques to fine-tune large language models (LLM). A weakn…

Diversity

Charting Empirical Laws for LLM Fine-Tuning in Scientific Multi-Discipline Learning

2026-02-11 · Lintao Wang, Zhuqiang Lu, Yilin Zhu, Kun Hu 외 arxiv

While large language models (LLMs) have achieved strong performance through fine-tuning within individual scientific domains, their learning dynamics in multi-disciplinary contexts remains poorly understood, despite the …

Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs

2025-09-25 · Honglin Zhang, Qianyue Hao, Fengli Xu, Yong Li arxiv

Large language models (LLMs) acquire extensive prior knowledge through large-scale pretraining and can be further enhanced via supervised fine-tuning (SFT) or reinforcement learning (RL)-based post-training. A growing bo…

Reinforcement Learning

Learning Loss Landscapes in Preference Optimization

2024-11-10 · Carlo Alfano, Silvia Sapora, Jakob Nicolaus Foerster, Patrick Rebeschini 외

We present an empirical study investigating how specific properties of preference datasets, such as mixed-quality or noisy data, affect the performance of Preference Optimization (PO) algorithms. Our experiments, conduct…

MuJoCo

Evaluating the Diversity and Quality of LLM Generated Content

2025-04-16 · Alexander Shypula, Shuo Li, Botong Zhang, Vishakh Padmakumar 외

Recent work suggests that preference-tuning techniques--including Reinforcement Learning from Human Preferences (RLHF) methods like PPO and GRPO, as well as alternatives like DPO--reduce diversity, creating a dilemma giv…

DiversitySynthetic Data Generation