paper-with-me

Papers

A Taxonomy of Stereotype Content in Large Language Models

2024-07-31 · Gandalf Nicolas, Aylin Caliskan

This study introduces a taxonomy of stereotype content in contemporary large language models (LLMs). We prompt ChatGPT 3.5, Llama 3, and Mixtral 8x7B, three powerful and widely used LLMs, for the characteristics associated with 87 social categories (e.g., gender, race, occupations). We identify 14 stereotype dimensions (e.g., Morality, Ability, Health, Beliefs, Emotions), accounting for ~90% of LLM stereotype associations. Warmth and Competence facets were the most frequent content, but all other dimensions were significantly prevalent. Stereotypes were more positive in LLMs (vs. humans), but there was significant variability across categories and dimensions. Finally, the taxonomy predicted the LLMs' internal evaluations of social categories (e.g., how positively/negatively the categories were represented), supporting the relevance of a multidimensional taxonomy for characterizing LLM stereotypes. Our findings suggest that high-dimensional human stereotypes are reflected in LLMs and must be considered in AI auditing and debiasing to minimize unidentified harms from reliance in low-dimensional views of bias in LLMs.

📄 PDF Abstract BibTeX arXiv:2408.00162

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

Incorporating Human Explanations for Robust Hate Speech Detection

2024-11-09 · Jennifer L. Chen, Faisal Ladhak, Daniel Li, Noémie Elhadad

Given the black-box nature and complexity of large transformer language models (LM), concerns about generalizability and robustness present ethical implications for domains such as hate speech (HS) detection. Using the c…

Hate Speech Detection

SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models

2023-05-19 · Akshita Jha, Aida Davani, Chandan K. Reddy, Shachi Dave 외

Stereotype benchmark datasets are crucial to detect and mitigate social stereotypes about groups of people in NLP models. However, existing datasets are limited in size and coverage, and are largely restricted to stereot…

A Robust Bias Mitigation Procedure Based on the Stereotype Content Model

2022-10-26 · Eddie L. Ungless, Amy Rafferty, Hrichika Nag, Björn Ross

The Stereotype Content model (SCM) states that we tend to perceive minority groups as cold, incompetent or both. In this paper we adapt existing work to demonstrate that the Stereotype Content model holds for contextuali…

Language ModelingLanguage ModellingWord Embeddings

Identifying Implicit Social Biases in Vision-Language Models

2024-11-01 · Kimia Hamidieh, Haoran Zhang, Walter Gerych, Thomas Hartvigsen 외

Vision-language models, like CLIP (Contrastive Language Image Pretraining), are becoming increasingly popular for a wide range of multimodal retrieval tasks. However, prior work has shown that large language and deep vis…

Fairness

Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech

2026-09-15 · Greta Damo, Elias Urios Alacreu, Elena Cabrio, Paolo Rosso 외 arxiv

Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, …