paper-with-me

홈 › Papers

From 1,000,000 Users to Every User: Scaling Up Personalized Preference for User-level Alignment

2025-03-19 · Jia-Nan Li, Jian Guan, Songhao Wu, Wei Wu, Rui Yan

Large language models (LLMs) have traditionally been aligned through one-size-fits-all approaches that assume uniform human preferences, fundamentally overlooking the diversity in user values and needs. This paper introduces a comprehensive framework for scalable personalized alignment of LLMs. We establish a systematic preference space characterizing psychological and behavioral dimensions, alongside diverse persona representations for robust preference inference in real-world scenarios. Building upon this foundation, we introduce \textsc{AlignX}, a large-scale dataset of over 1.3 million personalized preference examples, and develop two complementary alignment approaches: \textit{in-context alignment} directly conditioning on persona representations and \textit{preference-bridged alignment} modeling intermediate preference distributions. Extensive experiments demonstrate substantial improvements over existing methods, with an average 17.06\% accuracy gain across four benchmarks while exhibiting a strong adaptation capability to novel preferences, robustness to limited user data, and precise preference controllability. These results validate our framework's effectiveness, advancing toward truly user-adaptive AI systems.

📄 PDF Abstract BibTeX arXiv:2503.15463

Code (1)

jinaleejnl/alignx 공식 구현 pytorch

Tasks

Diversity

Similar Papers 제목 키워드 기반

P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling

2026-02-12 · Pinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen 외 arxiv

Personalized alignment of large language models seeks to adapt responses to individual user preferences, typically via reinforcement learning. A key challenge is obtaining accurate, user-specific reward signals in open-e…

Reinforcement Learning

Learning to summarize user information for personalized reinforcement learning from human feedback

2025-07-17 · Hyunji Nam, Yanming Wan, Mickel Liu, Peter Ahnn 외 arxiv

As everyday use cases of large language model (LLM) AI assistants have expanded, it is becoming increasingly important to personalize responses to align to different users' preferences and goals. While reinforcement lear…

Reinforcement Learning

Efficient Visual Appearance Optimization by Learning from Prior Preferences

2025-07-21 · Zhipeng Li, Yi-Chi Liao, Christian Holz arxiv

Adjusting visual parameters such as brightness and contrast is common in our everyday experiences. Finding the optimal parameter setting is challenging due to the large search space and the lack of an explicit objective …

Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs

2025-02-26 · Zhaowei Zhang, Fengshuo Bai, Qizhi Chen, Chengdong Ma 외

How to align large language models (LLMs) with user preferences from a static general dataset has been frequently studied. However, user preferences are usually personalized, changing, and diverse regarding culture, valu…

Computational Efficiency

CoBaR: Confidence-Based Recommender

2018-08-21 · Fernando S. Aguiar Neto, Arthur F. da Costa, Marcelo G. Manzato

Neighborhood-based collaborative filtering algorithms usually adopt a fixed neighborhood size for every user or item, although groups of users or items may have different lengths depending on users' preferences. In this …

ClusteringCollaborative Filtering