paper-with-me

홈 › Papers

MIST: Towards Multi-dimensional Implicit Bias and Stereotype Evaluation of LLMs via Theory of Mind

2025-06-17 · Yanlin Li, Hao liu, Huimin Liu, Yinwei Wei, Yupeng Hu

Theory of Mind (ToM) in Large Language Models (LLMs) refers to their capacity for reasoning about mental states, yet failures in this capacity often manifest as systematic implicit bias. Evaluating this bias is challenging, as conventional direct-query methods are susceptible to social desirability effects and fail to capture its subtle, multi-dimensional nature. To this end, we propose an evaluation framework that leverages the Stereotype Content Model (SCM) to reconceptualize bias as a multi-dimensional failure in ToM across Competence, Sociability, and Morality. The framework introduces two indirect tasks: the Word Association Bias Test (WABT) to assess implicit lexical associations and the Affective Attribution Test (AAT) to measure covert affective leanings, both designed to probe latent stereotypes without triggering model avoidance. Extensive experiments on 8 State-of-the-Art LLMs demonstrate our framework's capacity to reveal complex bias structures, including pervasive sociability bias, multi-dimensional divergence, and asymmetric stereotype amplification, thereby providing a more robust methodology for identifying the structural nature of implicit bias.

📄 PDF Abstract BibTeX arXiv:2506.14161

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models

2025-10-09 · Konrad Löhr, Shuzhou Yuan, Michael Färber arxiv

Large Language Models (LLMs) are increasingly integral to information dissemination and decision-making processes. Given their growing societal influence, understanding potential biases, particularly within the political…

FairMonitor: A Four-Stage Automatic Framework for Detecting Stereotypes and Biases in Large Language Models

2023-08-21 · Yanhong Bai, Jiabao Zhao, Jinxin Shi, Tingjiang Wei 외

Detecting stereotypes and biases in Large Language Models (LLMs) can enhance fairness and reduce adverse impacts on individuals or groups when these LLMs are applied. However, the majority of existing methods focus on me…

Fairness

Seeing Stereotypes

2025-03-04 · Elisa Baldazzi, Pietro Biroli, Marina Della Giusta, Florent Dubois

Reliance on stereotypes is a persistent feature of human decision-making and has been extensively documented in educational settings, where it can shape students' confidence, performance, and long-term human capital accu…

Decision MakingSurveyvalid

Probing Explicit and Implicit Gender Bias through LLM Conditional Text Generation

2023-11-01 · Xiangjue Dong, Yibo Wang, Philip S. Yu, James Caverlee

Large Language Models (LLMs) can generate biased and toxic responses. Yet most prior work on LLM gender bias evaluation requires predefined gender-related phrases or gender stereotypes, which are challenging to be compre…

Conditional Text GenerationFairnessText Generation

FairMonitor: A Dual-framework for Detecting Stereotypes and Biases in Large Language Models

2024-05-06 · Yanhong Bai, Jiabao Zhao, Jinxin Shi, Zhentao Xie 외

Detecting stereotypes and biases in Large Language Models (LLMs) is crucial for enhancing fairness and reducing adverse impacts on individuals or groups when these models are applied. Traditional methods, which rely on e…

Fairness