paper-with-me

홈 › Papers

Underrepresentation, Label Bias, and Proxies: Towards Data Bias Profiles for the EU AI Act and Beyond

2025-07-09 · Marina Ceccon, Giandomenico Cornacchia, Davide Dalle Pezze, Alessandro Fabris, Gian Antonio Susto arxiv

Undesirable biases encoded in the data are key drivers of algorithmic discrimination. Their importance is widely recognized in the algorithmic fairness literature, as well as legislation and standards on anti-discrimination in AI. Despite this recognition, data biases remain understudied, hindering the development of computational best practices for their detection and mitigation. In this work, we present three common data biases and study their individual and joint effect on algorithmic discrimination across a variety of datasets, models, and fairness measures. We find that underrepresentation of vulnerable populations in training sets is less conducive to discrimination than conventionally affirmed, while combinations of proxies and label bias can be far more critical. Consequently, we develop dedicated mechanisms to detect specific types of bias, and combine them into a preliminary construct we refer to as the Data Bias Profile (DBP). This initial formulation serves as a proof of concept for how different bias signals can be systematically documented. Through a case study with popular fairness datasets, we demonstrate the effectiveness of the DBP in predicting the risk of discriminatory outcomes and the utility of fairness-enhancing interventions. Overall, this article bridges algorithmic fairness research and anti-discrimination policy through a data-centric lens.

📄 PDF Abstract BibTeX arXiv:2507.08866

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Correcting Underrepresentation and Intersectional Bias for Classification

2023-06-19 · Emily Diana, Alexander Williams Tolbert

We consider the problem of learning from data corrupted by underrepresentation bias, where positive examples are filtered from the data at different, unknown rates for a fixed number of sensitive groups. We show that wit…

Classification

Shedding light on underrepresentation and Sampling Bias in machine learning

2023-06-08 · Sami Zhioua, Rūta Binkytė

Accurately measuring discrimination is crucial to faithfully assessing fairness of trained machine learning (ML) models. Any bias in measuring discrimination leads to either amplification or underestimation of the existi…

Fairness

Distributionally Robust Optimization and Invariant Representation Learning for Addressing Subgroup Underrepresentation: Mechanisms and Limitations

2023-08-12 · Nilesh Kumar, Ruby Shrestha, Zhiyuan Li, Linwei Wang

Spurious correlation caused by subgroup underrepresentation has received increasing attention as a source of bias that can be perpetuated by deep neural networks (DNNs). Distributionally robust optimization has shown suc…

image-classificationImage ClassificationMedical Image ClassificationRepresentation Learning

Bayesian generative models can flag performance loss, bias, and out-of-distribution image content

2025-03-21 · Miguel López-Pérez, Marco Miani, Valery Naranjo, Søren Hauberg 외

Generative models are popular for medical imaging tasks such as anomaly detection, feature extraction, data visualization, or image generation. Since they are parameterized by deep learning models, they are often sensiti…

Anomaly DetectionData VisualizationImage GenerationUncertainty Quantification

Do LLMs exhibit human-like response biases? A case study in survey design

2023-11-07 · Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar 외

As large language models (LLMs) become more capable, there is growing excitement about the possibility of using LLMs as proxies for humans in real-world tasks where subjective labels are desired, such as in surveys and o…

Survey