paper-with-me

홈 › Papers

Benchmarking and Improving LLM Robustness for Personalized Generation

2025-09-18 · Chimaobi Okite, Naihao Deng, Kiran Bodipati, Huaidian Hou, Joyce Chai, Rada Mihalcea arxiv

Recent years have witnessed a growing interest in personalizing the responses of large language models (LLMs). While existing evaluations primarily focus on whether a response aligns with a user's preferences, we argue that factuality is an equally important yet often overlooked dimension. In the context of personalization, we define a model as robust if its responses are both factually accurate and align with the user preferences. To assess this, we introduce PERG, a scalable framework for evaluating robustness in LLMs, along with a new dataset, PERGData. We evaluate fourteen models from five different model families using different prompting methods. Our findings show that current LLMs struggle with robust personalization: even the strongest models (GPT-4.1, LLaMA3-70B) fail to maintain correctness in 5% of previously successful cases without personalization, while smaller models (e.g., 7B-scale) can fail more than 20% of the time. Further analysis reveals that robustness is significantly affected by the nature of the query and the type of user preference. To mitigate these failures, we propose Pref-Aligner, a two-stage approach that improves robustness by an average of 25% across models. Our work highlights critical gaps in current evaluation practices and introduces tools and metrics to support more reliable, user-aligned LLM deployments.

📄 PDF Abstract BibTeX arXiv:2509.19358

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Personalized Benchmarking with the Ludwig Benchmarking Toolkit

2021-11-08 · Avanika Narayan, Piero Molino, Karan Goel, Willie Neiswanger 외

The rapid proliferation of machine learning models across domains and deployment settings has given rise to various communities (e.g. industry practitioners) which seek to benchmark models across tasks and objectives of …

BenchmarkingHyperparameter Optimizationtext-classificationText Classification

DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

2024-06-24 · Yuang Peng, Yuxin Cui, Haomiao Tang, Zekun Qi 외

Personalized image generation holds great promise in assisting humans in everyday work and life due to its impressive function in creatively generating personalized content. However, current evaluations either are automa…

BenchmarkingImage GenerationPersonalized Image Generation

PersoBench: Benchmarking Personalized Response Generation in Large Language Models

2024-10-04 · Saleh Afzoon, Usman Naseem, Amin Beheshti, Zahra Jamali

While large language models (LLMs) have exhibited impressive conversational capabilities, their proficiency in delivering personalized responses remains unclear. Although recent benchmarks automatically evaluate persona …

BenchmarkingDialogue GenerationDiversityResponse Generation

PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models

2026-01-07 · Yuwen Wang, Xinyuan Qian, Tian-Hao Zhang, Jiaran Gao 외 arxiv

Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is largely generic (e.g., summarizing spoken con…

Question Answering

RetailSynth: Synthetic Data Generation for Retail AI Systems Evaluation

2023-12-21 · Yu Xia, Ali Arian, Sriram Narayanamoorthy, Joshua Mabry

Significant research effort has been devoted in recent years to developing personalized pricing, promotions, and product recommendation algorithms that can leverage rich customer data to learn and earn. Systematic benchm…

BenchmarkingProduct RecommendationSensitivitySynthetic Data Generation