paper-with-me

Papers

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

2026-05-11 · Pierre Le Jeune, Étienne Duchesne, Weixuan Xiao, Stefano Palminteri, Bazire Houssin, Benoît Malézieux, Matteo Dora arxiv

Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual dataset and evaluation pipeline for systematically studying the emergence of social bias in open-ended LLM generation. The dataset covers 10 languages and 79 socio-demographic attributes, and comprises over 650k stories generated by 23 recent LLMs, each annotated with the socio-demographic profile of the protagonist across 19 dimensions. From these, we apply statistical tests to identify more than 1{,}500 over-represented associations, which we then rate for harmfulness through both a panel of humans (N = 247) and the same LLMs. We report three main findings. \textbf{(i)} Every model we evaluate emits consequential harmful stereotypes in open-ended generation, regardless of size or capabilities, and these associations are largely shared across providers rather than isolated misbehaviors. \textbf{(ii)} Prompt language strongly shapes which stereotypes appear: rather than transferring as a shared set of biases, harmful associations adapt culturally to the prompt language and amplify bias against locally salient protected groups. \textbf{(iii)} Human and LLM harmfulness judgments are broadly aligned (Spearman $ρ=0.62$), with disagreements concentrating on specific attribute classes rather than specific providers. To support further analyses, we release the evaluation code and the dataset, including model generations, attribute annotations, and harmfulness ratings.

📄 PDF Abstract BibTeX arXiv:2605.10442

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs

2025-11-03 · Muhammed Saeed, Muhammad Abdul-mageed, Shady Shehata arxiv

Large language models (LLMs) are widely deployed for open-ended communication, yet most bias evaluations still rely on English, classification-style tasks. We introduce \corpusname, a new multilingual, debate-style bench…

MBBQ: A Dataset for Cross-Lingual Comparison of Stereotypes in Generative LLMs

2024-06-11 · Vera Neplenbroek, Arianna Bisazza, Raquel Fernández

Generative large language models (LLMs) have been shown to exhibit harmful biases and stereotypes. While safety fine-tuning typically takes place in English, if at all, these models are being used by speakers of many dif…

Question Answering

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

2026-07-23 · Emilio Ferrara arxiv

Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its prin…

Multilingual large language models leak human stereotypes across language boundaries

2023-12-12 · Yang Trista Cao, Anna Sotnikova, Jieyu Zhao, Linda X. Zou 외

Multilingual large language models have gained prominence for their proficiency in processing and generating text across languages. Like their monolingual counterparts, multilingual models are likely to pick up on stereo…

CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

2025-05-26 · Ruixiang Feng, Shen Gao, Xiuying Chen, Lisi Chen 외

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they often exhibit a specific cultural biases, neglecting the values and linguistic diversity of low-resource regions. This…

DiversityOpen-Ended Question AnsweringQuestion Answering