paper-with-me

홈 › Papers

Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study

2026-02-19 · Kensuke Okada, Yui Furukawa, Kyosuke Bunji arxiv

Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safety and bias assessments. Yet these instruments presume honest responding; in evaluative contexts, LLMs can instead gravitate toward socially preferred answers-a form of socially desirable responding (SDR)-biasing questionnaire-derived scores and downstream conclusions. We propose a psychometric framework to quantify and mitigate SDR in questionnaire-based evaluation of LLMs. To quantify SDR, the same inventory is administered under HONEST versus FAKE-GOOD instructions, and SDR is computed as a direction-corrected standardized effect size from item response theory (IRT)-estimated latent scores. This enables comparisons across constructs and response formats, as well as against human instructed-faking benchmarks. For mitigation, we construct a graded forced-choice (GFC) Big Five inventory by selecting 30 cross-domain pairs from an item pool via constrained optimization to match desirability. Across nine instruction-following LLMs evaluated on synthetic personas with known target profiles, Likert-style questionnaires show consistently large SDR, whereas desirability-matched GFC substantially attenuates SDR while largely preserving the recovery of the intended persona profiles. These results highlight a model-dependent SDR-recovery trade-off and motivate SDR-aware reporting practices for questionnaire-based benchmarking and auditing of LLMs.

📄 PDF Abstract BibTeX arXiv:2602.17262

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Performance: Quantifying and Mitigating Label Bias in LLMs

2024-05-04 · Yuval Reif, Roy Schwartz

Large language models (LLMs) have shown remarkable adaptability to diverse tasks, by leveraging context prompts containing instructions, or minimal input-output examples. However, recent work revealed they also exhibit l…

LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models

2025-05-21 · Zhanyue Qin, Yue Ding, Deyuan Liu, Qingbin Liu 외

Nowadays, Large Language Models (LLMs) have attracted widespread attention due to their powerful performance. However, due to the unavoidable exposure to socially biased data during training, LLMs tend to exhibit social …

Misaligned by Reward: Socially Undesirable Preferences in LLMs

2026-05-06 · Gayane Ghazaryan, Esra Dönmez arxiv

Reward models are a key component of large language model alignment, serving as proxies for human preferences during training. However, existing evaluations focus primarily on broad instruction-following benchmarks, prov…

"A Woman is More Culturally Knowledgeable than A Man?": The Effect of Personas on Cultural Norm Interpretation in LLMs

2024-09-18 · Mahammed Kamruzzaman, Hieu Nguyen, Nazmul Hassan, Gene Louis Kim

As the deployment of large language models (LLMs) expands, there is an increasing demand for personalized LLMs. One method to personalize and guide the outputs of these models is by assigning a persona -- a role that des…

Mitigating Social Desirability Bias in Random Silicon Sampling

2025-12-27 · Sashank Chapala, Maksym Mironov, Songgaojun Deng arxiv

Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias …