Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
We present a large-scale study of gender bias in occupation classification, a task where the use of machine learning may lead to negative outcomes on peoples' lives. We analyze the potential allocation harms that can result from semantic representation bias. To do so, we study the impact on occupation classification of including explicit gender indicators---such as first names and pronouns---in different semantic representations of online biographies. Additionally, we quantify the bias that remains when these indicators are "scrubbed," and describe proxy behavior that occurs in the absence of explicit gender indicators. As we demonstrate, differences in true positive rates between genders are correlated with existing gender imbalances in occupations, which may compound these imbalances.
Code (5)
Tasks
ClassificationGeneral ClassificationSimilar Papers 제목 키워드 기반
Choosing the Better Bandit Algorithm under Data Sharing: When Do A/B Experiments Work?
We study A/B experiments that are designed to compare the performance of two recommendation algorithms. Prior work has shown that the standard difference-in-means estimator is biased in estimating the global treatment ef…
An Empirical Analysis of Fine-Tuning Large Language Models on Bioinformatics Literature: PRSGPT and BioStarsGPT
Large language models (LLMs) often lack specialized knowledge for complex bioinformatics applications. We present a reproducible pipeline for fine-tuning LLMs on specialized bioinformatics data, demonstrated through two …
parameter-efficient fine-tuningNatural Language InferenceDecoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
Understanding and mitigating the potential risks associated with foundation models (FMs) hinges on developing effective interpretability methods. Sparse Autoencoders (SAEs) have emerged as a promising tool for disentangl…
Interpretable machine learning applied to on-farm biosecurity and porcine reproductive and respiratory syndrome virus
Effective biosecurity practices in swine production are key in preventing the introduction and dissemination of infectious pathogens. Ideally, biosecurity practices should be chosen by their impact on bio-containment and…
BenchmarkingBIG-bench Machine LearningInterpretable Machine LearningContext vs Target Word: Quantifying Biases When Applying Models to Lexical Semantic Datasets
State-of-the-art contextualized models such as BERT use tasks such as WiC and WSD to evaluate their word-in-context representations. This inherently assumes that performance in these tasks reflect how well a model repres…
Entity Linking