paper-with-me

Papers

Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

2024-02-29 · Yiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding, Zexu Sun, Bowen Sun, Huimin Chen, Ruobing Xie, Jie zhou, Yankai Lin, Zhiyuan Liu, Maosong Sun

Alignment in artificial intelligence pursues the consistency between model responses and human preferences as well as values. In practice, the multifaceted nature of human preferences inadvertently introduces what is known as the "alignment tax" -a compromise where enhancements in alignment within one objective (e.g.,harmlessness) can diminish performance in others (e.g.,helpfulness). However, existing alignment techniques are mostly unidirectional, leading to suboptimal trade-offs and poor flexibility over various objectives. To navigate this challenge, we argue the prominence of grounding LLMs with evident preferences. We introduce controllable preference optimization (CPO), which explicitly specifies preference scores for different objectives, thereby guiding the model to generate responses that meet the requirements. Our experimental analysis reveals that the aligned models can provide responses that match various preferences among the "3H" (helpfulness, honesty, harmlessness) desiderata. Furthermore, by introducing diverse data and alignment goals, we surpass baseline methods in aligning with single objectives, hence mitigating the impact of the alignment tax and achieving improvements in multi-objective alignment.

📄 PDF Abstract BibTeX arXiv:2402.19085

Code (1)

OpenBMB/CPO 공식 구현 pytorch

Tasks

Navigate

Similar Papers 제목 키워드 기반

Controllable Expensive Multi-objective Learning with Warm-starting Bayesian Optimization

2023-11-26 · Quang-Huy Nguyen, Long P. Hoang, Hoang V. Viet, Dung D. Le

Pareto Set Learning (PSL) is a promising approach for approximating the entire Pareto front in multi-objective optimization (MOO) problems. However, existing derivative-free PSL methods are often unstable and inefficient…

Bayesian OptimizationGaussian Processes

MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models

2026-05-25 · Linhao Luo, Thuy-Trang Vu, Van-Anh Nguyen, Junae Kim 외 arxiv

Aligning large language models (LLMs) with diverse and multifaceted user preferences is a fundamental challenge in personalized AI systems. Existing multi-objective alignment methods either rely on costly training or req…

Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment

2024-08-12 · Karel D'Oosterlinck, Winnie Xu, Chris Develder, Thomas Demeester 외

Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes alignment a complicated procedure, somet…

Contrastive Learning

Controllable Pareto Trade-off between Fairness and Accuracy

2025-09-17 · Yongkang Du, Jieyu Zhao, Yijun Yang, Tianyi Zhou arxiv

The fairness-accuracy trade-off is a key challenge in NLP tasks. Current work focuses on finding a single "optimal" solution to balance the two objectives, which is limited considering the diverse solutions on the Pareto…

Hate Speech Detection

Controllable Pareto Multi-Task Learning

2020-10-13 · Xi Lin, Zhiyuan Yang, Qingfu Zhang, Sam Kwong

A multi-task learning (MTL) system aims at solving multiple related tasks at the same time. With a fixed model capacity, the tasks would be conflicted with each other, and the system usually has to make a trade-off among…

Multiobjective OptimizationMulti-Task Learning