paper-with-me

홈 › Papers

Mitigating Social Desirability Bias in Random Silicon Sampling

2025-12-27 · Sashank Chapala, Maksym Mironov, Songgaojun Deng arxiv

Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data toward socially acceptable answers. Existing studies on social desirability bias in LLM-based sampling remain limited. In this work, we investigate whether minimal, psychologically grounded prompt wording can mitigate this bias and improve alignment between silicon and human samples. We conducted a study using data from the American National Election Study (ANES) on three LLMs from two model families: the open-source Llama-3.1 series and GPT-4.1-mini. We first replicate a baseline silicon sampling study, confirming the persistent Social Desirability Bias. We then test four prompt-based mitigation methods: \emph{reformulated} (neutral, third-person phrasing), \emph{reverse-coded} (semantic inversion), and two meta-instructions, \emph{priming} and \emph{preamble}, respectively encouraging analytics and sincerity. Alignment with ANES is evaluated using Jensen-Shannon Divergence with bootstrap confidence intervals. Our results demonstrate that reformulated prompts most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES. Reverse-coding produced mixed results across eligible items, while the Priming and Preamble encouraged response uniformity and showed no systematic benefit for bias mitigation. Our findings validate the efficacy of prompt-based framing controls in mitigating inherent Social Desirability Bias in LLMs, providing a practical path toward more representative silicon samples.

📄 PDF Abstract BibTeX arXiv:2512.22725

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Large Language Models Show Human-like Social Desirability Biases in Survey Responses

2024-05-09 · Aadesh Salecha, Molly E. Ireland, Shashanka Subrahmanya, João Sedoc 외

As Large Language Models (LLMs) become widely used to model and simulate human behavior, understanding their biases becomes critical. We developed an experimental framework using Big Five personality surveys and uncovere…

Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias

2025-09-09 · Sirui Wu, Daijin Yang arxiv

This study evaluates item neutralization assisted by the large language model (LLM) to reduce social desirability bias in personality assessment. GPT-o3 was used to rewrite the International Personality Item Pool Big Fiv…

Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study

2026-02-19 · Kensuke Okada, Yui Furukawa, Kyosuke Bunji arxiv

Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safety and bias assessments. Yet these instruments presume honest responding;…

Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations

2024-10-20 · Sanguk Lee, Kai-Qi Yang, Tai-Quan Peng, Ruth Heo 외

Large language models (LLMs) are employed to simulate human-like responses in social surveys, yet it remains unclear if they develop biases like social desirability response (SDR) bias. To investigate this, GPT-4 was ass…

The Narcissus Hypothesis: Descending to the Rung of Illusion

2025-09-22 · Riccardo Cadei, Christian Internò arxiv

Modern foundational models increasingly reflect not just world knowledge, but patterns of human preference embedded in their training data. We hypothesize that recursive alignment-via human feedback and model-generated c…