Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets
Auditing NLP systems for computational harms like surfacing stereotypes is an elusive goal. Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system{'}s behavior on these pairs into measurements of harms. We examine four such benchmarks constructed for two NLP tasks: language modeling and coreference resolution. We apply a measurement modeling lens{---}originating from the social sciences{---}to inventory a range of pitfalls that threaten these benchmarks{'} validity as measurement models for stereotyping. We find that these benchmarks frequently lack clear articulations of what is being measured, and we highlight a range of ambiguities and unstated assumptions that affect how these benchmarks conceptualize and operationalize stereotyping.
Code (0)
등록된 구현이 없습니다.
Tasks
coreference-resolutionCoreference ResolutionFairnessLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Fairness in representation: quantifying stereotyping as a representational harm
While harms of allocation have been increasingly studied as part of the subfield of algorithmic fairness, harms of representation have received considerably less attention. In this paper, we formalize two notions of ster…
BIG-bench Machine LearningFairnessWait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deducti…
Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for English-language use (Claude, GPT, Gemini) and three for East Asian use …
FairPrism: Evaluating Fairness-Related Harms in Text Generation
It is critical to measure and mitigate fairness- related harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-gene…
FairnessText GenerationLidar-based Norwegian tree species detection using deep learning
Background: The mapping of tree species within Norwegian forests is a time-consuming process, involving forest associations relying on manual labeling by experts. The process can involve both aerial imagery, personal fam…
Deep LearningSemantic Segmentation