paper-with-me

홈 › Papers

Dataset Representativeness and Downstream Task Fairness

2024-06-28 · Victor Borza, Andrew Estornell, Chien-Ju Ho, Bradley Malin, Yevgeniy Vorobeychik

Our society collects data on people for a wide range of applications, from building a census for policy evaluation to running meaningful clinical trials. To collect data, we typically sample individuals with the goal of accurately representing a population of interest. However, current sampling processes often collect data opportunistically from data sources, which can lead to datasets that are biased and not representative, i.e., the collected dataset does not accurately reflect the distribution of demographics of the true population. This is a concern because subgroups within the population can be under- or over-represented in a dataset, which may harm generalizability and lead to an unequal distribution of benefits and harms from downstream tasks that use such datasets (e.g., algorithmic bias in medical decision-making algorithms). In this paper, we assess the relationship between dataset representativeness and group-fairness of classifiers trained on that dataset. We demonstrate that there is a natural tension between dataset representativeness and classifier fairness; empirically we observe that training datasets with better representativeness can frequently result in classifiers with higher rates of unfairness. We provide some intuition as to why this occurs via a set of theoretical results in the case of univariate classifiers. We also find that over-sampling underrepresented groups can result in classifiers which exhibit greater bias to those groups. Lastly, we observe that fairness-aware sampling strategies (i.e., those which are specifically designed to select data with high downstream fairness) will often over-sample members of majority groups. These results demonstrate that the relationship between dataset representativeness and downstream classifier fairness is complex; balancing these two quantities requires special care from both model- and dataset-designers.

📄 PDF Abstract BibTeX arXiv:2407.00170

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

FAL-CUR: Fair Active Learning using Uncertainty and Representativeness on Fair Clustering

2022-09-21 · Ricky Fajri, Akrati Saxena, Yulong Pei, Mykola Pechenizkiy

Active Learning (AL) techniques have proven to be highly effective in reducing data labeling costs across a range of machine learning tasks. Nevertheless, one known challenge of these methods is their potential to introd…

Active LearningClusteringFairness

Balanced Ranking with Diversity Constraints

2019-06-04 · Ke Yang, Vasilis Gkatzelis, Julia Stoyanovich

Many set selection and ranking algorithms have recently been enhanced with diversity constraints that aim to explicitly increase representation of historically disadvantaged populations, or to improve the overall represe…

DiversityFairness

RS-Prune: Training-Free Data Pruning at High Ratios for Efficient Remote Sensing Diffusion Foundation Models

2025-12-29 · Fan Wei, Runmin Dong, Yushan Lai, Yixiang Yang 외 arxiv

Diffusion-based remote sensing (RS) generative foundation models are cruial for downstream tasks. However, these models rely on large amounts of globally representative data, which often contain redundancy, noise, and cl…

Scene Classification

Artificial Delegates Resolve Fairness Issues in Perpetual Voting with Partial Turnout

2025-06-26 · Apurva Shah, Axel Abels, Ann Nowé, Tom Lenaerts

Perpetual voting addresses fairness in sequential collective decision-making by evaluating representational equity over time. However, existing perpetual voting rules rely on full participation and complete approval info…

Decision MakingFairness

Learning Fair Representations with High-Confidence Guarantees

2023-10-23 · Yuhong Luo, Austin Hoag, Philip S. Thomas

Representation learning is increasingly employed to generate representations that are predictive across multiple downstream tasks. The development of representation learning algorithms that provide strong fairness guaran…

AllFairnessRepresentation Learning