paper-with-me

홈 › Papers

Evaluating Large Language Model Biases in Persona-Steered Generation

2024-05-30 · Andy Liu, Mona Diab, Daniel Fried

The task of persona-steered text generation requires large language models (LLMs) to generate text that reflects the distribution of views that an individual fitting a persona could have. People have multifaceted personas, but prior work on bias in LLM-generated opinions has only explored multiple-choice settings or one-dimensional personas. We define an incongruous persona as a persona with multiple traits where one trait makes its other traits less likely in human survey data, e.g. political liberals who support increased military spending. We find that LLMs are 9.7% less steerable towards incongruous personas than congruous ones, sometimes generating the stereotypical stance associated with its demographic rather than the target stance. Models that we evaluate that are fine-tuned with Reinforcement Learning from Human Feedback (RLHF) are more steerable, especially towards stances associated with political liberals and women, but present significantly less diverse views of personas. We also find variance in LLM steerability that cannot be predicted from multiple-choice opinion evaluation. Our results show the importance of evaluating models in open-ended text generation, as it can surface new LLM opinion biases. Moreover, such a setup can shed light on our ability to steer models toward a richer and more diverse range of viewpoints.

📄 PDF Abstract BibTeX arXiv:2405.20253

Code (1)

andyjliu/persona-steered-generation-bias 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultiple-choiceText Generation

Similar Papers 제목 키워드 기반

GermanPartiesQA: Benchmarking Commercial Large Language Models for Political Bias and Sycophancy

2024-07-25 · Jan Batzner, Volker Stocker, Stefan Schmid, Gjergji Kasneci

LLMs are changing the way humans create and interact with content, potentially affecting citizens' political opinions and voting decisions. As LLMs increasingly shape our digital information ecosystems, auditing to evalu…

Benchmarking

Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models

2025-04-11 · Yin Jou Huang, Rafik Hadfi

Self-report questionnaires have long been used to assess LLM personality traits, yet they fail to capture behavioral nuances due to biases and meta-knowledge contamination. This paper proposes a novel multi-observer fram…

Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue Systems

2023-10-08 · Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng 외

Recent advancements in Large Language Models empower them to follow freeform instructions, including imitating generic or specific demographic personas in conversations. We define generic personas to represent demographi…

Benchmarking

Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)

2024-10-25 · Leander Girrbach, Stephan Alaniz, Yiran Huang, Trevor Darrell 외

Pre-trained large language models (LLMs) have been reliably integrated with visual input for multimodal tasks. The widespread adoption of instruction-tuned image-to-text vision-language assistants (VLAs) like LLaVA and I…

AttributeImage to text

Evaluating Biases in Context-Dependent Health Questions

2024-03-07 · Sharon Levy, Tahilin Sanchez Karver, William D. Adler, Michelle R. Kaufman 외

Chat-based large language models have the opportunity to empower individuals lacking high-quality healthcare access to receive personalized information across a variety of topics. However, users may ask underspecified qu…

Language ModelingLanguage ModellingLarge Language Model