paper-with-me

홈 › Papers

Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization

2025-05-22 · Vera Neplenbroek, Arianna Bisazza, Raquel Fernández

Generative Large Language Models (LLMs) infer user's demographic information from subtle cues in the conversation -- a phenomenon called implicit personalization. Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups, even when no demographic information is explicitly provided. In this work, we systematically explore how LLMs respond to stereotypical cues using controlled synthetic conversations, by analyzing the models' latent user representations through both model internals and generated answers to targeted user questions. Our findings reveal that LLMs do infer demographic attributes based on these stereotypical signals, which for a number of groups even persists when the user explicitly identifies with a different demographic group. Finally, we show that this form of stereotype-driven implicit personalization can be effectively mitigated by intervening on the model's internal representations using a trained linear probe to steer them toward the explicitly stated identity. Our results highlight the need for greater transparency and control in how LLMs represent user identity.

📄 PDF Abstract BibTeX arXiv:2505.16467

Code (1)

veranep/implicit-personalization-stereotypes 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Criteria for the Annotation of Implicit Stereotypes

2022-06-01 · LREC 2022 6 · Wolfgang Schmeisser-Nieto, Montserrat Nofre, Mariona Taulé

The growth of social media has brought with it a massive channel for spreading and reinforcing stereotypes. This issue becomes critical when the affected targets are minority groups such as women, the LGBT+ community and…

Sentence

Seeing Stereotypes

2025-03-04 · Elisa Baldazzi, Pietro Biroli, Marina Della Giusta, Florent Dubois

Reliance on stereotypes is a persistent feature of human decision-making and has been extensively documented in educational settings, where it can shape students' confidence, performance, and long-term human capital accu…

Decision MakingSurveyvalid

FairMonitor: A Four-Stage Automatic Framework for Detecting Stereotypes and Biases in Large Language Models

2023-08-21 · Yanhong Bai, Jiabao Zhao, Jinxin Shi, Tingjiang Wei 외

Detecting stereotypes and biases in Large Language Models (LLMs) can enhance fairness and reduce adverse impacts on individuals or groups when these LLMs are applied. However, the majority of existing methods focus on me…

Fairness

Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech

2026-09-15 · Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso 외 arxiv

Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, …

Language is Scary when Over-Analyzed: Unpacking Implied Misogynistic Reasoning with Argumentation Theory-Driven Prompts

2024-09-04 · Arianna Muti, Federico Ruggeri, Khalid Al-Khatib, Alberto Barrón-Cedeño 외

We propose misogyny detection as an Argumentative Reasoning task and we investigate the capacity of large language models (LLMs) to understand the implicit reasoning used to convey misogyny in both Italian and English. T…