SWEPO: Simultaneous Weighted Preference Optimization for Group Contrastive Alignment
We introduce Simultaneous Weighted Preference Optimization (SWEPO), a novel extension of Direct Preference Optimization (DPO) designed to accommodate multiple dynamically chosen positive and negative responses for each query. SWEPO employs a weighted group contrastive loss, assigning weights to responses based on their deviation from the mean reward score. This approach effectively prioritizes responses that are significantly better or worse than the average, enhancing optimization. Our theoretical analysis demonstrates that simultaneously considering multiple preferences reduces alignment bias, resulting in more robust alignment. Additionally, we provide insights into the training dynamics of our loss function and a related function, InfoNCA. Empirical validation on the UltraFeedback dataset establishes SWEPO as state-of-the-art, with superior performance in downstream evaluations using the AlpacaEval dataset.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image diffusion models. However, existing methods largely reduce supervisio…
Text-to-Image GenerationReinforcement LearningImage EditingAPPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
Aligning large language models (LLMs) with diverse human preferences requires pluralistic alignment, where a single model must respect the values of multiple distinct groups simultaneously. In federated reinforcement lea…
Reinforcement LearningGenerative AI Empowered Semantic Feature Multiple Access (SFMA) Over Wireless Networks
This paper investigates a novel generative artificial intelligence (GAI) empowered multi-user semantic communication system called semantic feature multiple access (SFMA) for video transmission, which comprises a base st…
Semantic CommunicationVideo Frame InterpolationUser Grouping and Resource Allocation in Multiuser MIMO Systems under SWIPT
This paper considers a broadcast multiple-input multiple-output (MIMO) network with multiple users and simultaneous wireless information and power transfer (SWIPT). In this scenario, it is assumed that some users are abl…
SchedulingNo Preference Left Behind: Group Distributional Preference Optimization
Preferences within a group of people are not uniform but follow a distribution. While existing alignment methods like Direct Preference Optimization (DPO) attempt to steer models to reflect human preferences, they strugg…
DiversityLanguage ModelingLanguage Modelling