paper-with-me

홈 › Papers

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

2026-08-05 · Yushi Sun, Yanjie Zhang, Rui Sheng hf

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.

📄 PDF Abstract BibTeX arXiv:2608.04570

Code (3)

Tavish9/awesome-daily-AI-arxiv ★ 113
Valiant-Cat/hfpaper
arxivsub/arXivSub_daily_arxiv ★ 4

Similar Papers 제목 키워드 기반

Understanding the Role of User Profile in the Personalization of Large Language Models

2024-06-22 · Bin Wu, Zhengyan Shi, Hossein A. Rahmani, Varsha Ramineni 외

Utilizing user profiles to personalize Large Language Models (LLMs) has been shown to enhance the performance on a wide range of tasks. However, the precise role of user profiles and their effect mechanism on LLMs remain…

MIRAGE: Defending Long-Form RAG Against Misinformation Pollution

2026-07-06 · Saadeldine Eletter, Ruihong Zeng, Yuxia Wang, Maxim Panov 외 arxiv

Retrieval-Augmented Generation (RAG) improves factuality by grounding LLMs in external evidence, but real-world retrieval is often polluted: semantically relevant passages may contain subtle misinformation, misleading fr…

The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs

2025-10-10 · Xi Fang, Weijie Xu, Yuchong Zhang, Stephanie Eckman 외 arxiv

When an AI assistant remembers that Sarah is a single mother working two jobs, does it interpret her stress differently than if she were a wealthy executive? As personalized AI systems increasingly incorporate long-term …

Emotional Intelligence

Are Large Language Models In-Context Personalized Summarizers? Get an iCOPERNICUS Test Done!

2024-09-30 · Divya Patel, Pathik Patel, Ankush Chander, Sourish Dasgupta 외

Large Language Models (LLMs) have succeeded considerably in In-Context-Learning (ICL) based summarization. However, saliency is subject to the users' specific preference histories. Hence, we need reliable In-Context Pers…

In-Context Learning

Evaluating the Hidden Costs of Personalization in Large Language Models

2026-08-28 · Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan 외 hf

While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfac…