paper-with-me

홈 › Papers

Benchmarking Bias in Large Language Models during Role-Playing

2024-11-01 · Xinyue Li, Zhenpeng Chen, Jie M. Zhang, Yiling Lou, Tianlin Li, Weisong Sun, Yang Liu, Xuanzhe Liu

Large Language Models (LLMs) have become foundational in modern language-driven applications, profoundly influencing daily life. A critical technique in leveraging their potential is role-playing, where LLMs simulate diverse roles to enhance their real-world utility. However, while research has highlighted the presence of social biases in LLM outputs, it remains unclear whether and to what extent these biases emerge during role-playing scenarios. In this paper, we introduce BiasLens, a fairness testing framework designed to systematically expose biases in LLMs during role-playing. Our approach uses LLMs to generate 550 social roles across a comprehensive set of 11 demographic attributes, producing 33,000 role-specific questions targeting various forms of bias. These questions, spanning Yes/No, multiple-choice, and open-ended formats, are designed to prompt LLMs to adopt specific roles and respond accordingly. We employ a combination of rule-based and LLM-based strategies to identify biased responses, rigorously validated through human evaluation. Using the generated questions as the benchmark, we conduct extensive evaluations of six advanced LLMs released by OpenAI, Mistral AI, Meta, Alibaba, and DeepSeek. Our benchmark reveals 72,716 biased responses across the studied LLMs, with individual models yielding between 7,754 and 16,963 biased responses, underscoring the prevalence of bias in role-playing contexts. To support future research, we have publicly released the benchmark, along with all scripts and experimental results.

📄 PDF Abstract BibTeX arXiv:2411.00585

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingFairnessMultiple-choice

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration

2024-09-17 · Xin Guan, Ze Wang, Nathaniel Demchak, Saloni Gupta 외

The development of unbiased large language models is widely recognized as crucial, yet existing benchmarks fall short in detecting biases due to limited scope, contamination, and lack of a fairness baseline. SAGED(bias) …

BenchmarkingcounterfactualFairnessSentiment Analysis

VisoGender: A dataset for benchmarking gender bias in image-text pronoun resolution

2023-06-21 · NeurIPS 2023 11 · Siobhan Mackenzie Hall, Fernanda Gonçalves Abrantes, Hanwen Zhu, Grace Sodunke 외

We introduce VisoGender, a novel dataset for benchmarking gender bias in vision-language models. We focus on occupation-related biases within a hegemonic system of binary gender, inspired by Winograd and Winogender schem…

BenchmarkingRetrieval

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

2026-01-08 · Zhiwei Liu, Yupen Cao, Yuechen Jiang, Mohsinul Kabir 외 arxiv

Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-authored corpora, LLMs may inherit a range of human biases. Behavioral bia…

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

2026-08-30 · Shaghayegh Kolli, Sina Emami, Moreno D'Incà, Pouyan Nejadi 외 hf

Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations can underpin many observed forms of stere…

Benchmarking bias: Expanding clinical AI model card to incorporate bias reporting of social and non-social factors

2023-11-21 · Carolina A. M. Heming, Mohamed Abdalla, Shahram Mohanna, Monish Ahluwalia 외

Clinical AI model reporting cards should be expanded to incorporate a broad bias reporting of both social and non-social factors. Non-social factors consider the role of other factors, such as disease dependent, anatomic…

Benchmarking