paper-with-me

홈 › Papers

Language Model Personalization via Reward Factorization

2025-03-08 · Idan Shenfeld, Felix Faltings, Pulkit Agrawal, Aldo Pacchiano

Modern large language models (LLMs) are optimized for human-aligned responses using Reinforcement Learning from Human Feedback (RLHF). However, existing RLHF approaches assume a universal preference model and fail to account for individual user preferences, limiting their effectiveness in personalized applications. We introduce a framework that extends RLHF to enable user personalization by leveraging the assumption that user preferences lie in a low-dimensional space. Instead of training a separate model per user, we represent user-specific rewards as a linear combination of base reward functions. Using only ~10 user responses, our method can infer user-specific rewards and align LLM outputs accordingly. We validate our approach through experiments with both synthetic and real users, demonstrating significant personalization achieved by our method. In human evaluations, our method achieves a 67% win rate over default GPT-4o responses.

📄 PDF Abstract BibTeX arXiv:2503.06358

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Uncertainty-Aware Variational Reward Factorization via Probabilistic Preference Bases for LLM Personalization

2026-04-01 · Gyuseok Lee, Wonbin Kweon, Zhenrui Yue, SeongKu Kang 외 arxiv

Reward factorization personalizes large language models (LLMs) by decomposing rewards into shared basis functions and user-specific weights. Yet, existing methods estimate user weights from scarce data in isolation and a…

Personalized Response Generation with Tensor Factorization

2021-08-01 · ACL (GEM) 2021 8 · Zhenghui Wang, Lingxiao Luo, Diyi Yang

Personalized response generation is essential for more human-like conversations. However, how to model user personalization information with no explicit user persona descriptions or demographics still remains under-inves…

DecoderLanguage ModelingLanguage ModellingResponse Generation

Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning

2025-10-21 · Chenghao Zhu, Meiling Tao, Tiannan Wang, Dongyi Ding 외 arxiv

Faithfully personalizing large language models (LLMs) to align with individual user preferences is a critical but challenging task. While supervised fine-tuning (SFT) quickly reaches a performance plateau, standard reinf…

Reinforcement Learning

PrefReward: Learning User Preference Matrix for Personalized Text Generation

2026-07-23 · Yue Wu, Chengbing Wang, Yimeng Bai, Xiaoyan Zhao 외 arxiv

Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. However, most existing personalization approaches rely on implicit re…

Text Generation

Learning from Natural Language Feedback for Personalized Question Answering

2025-08-14 · Alireza Salemi, Hamed Zamani arxiv

Personalization is crucial for enhancing both the effectiveness and user satisfaction of language technologies, particularly in information-seeking tasks like question answering. Current approaches for personalizing larg…

Reinforcement LearningResponse GenerationQuestion Answering