paper-with-me

홈 › Papers

The Dataset Multiplicity Problem: How Unreliable Data Impacts Predictions

2023-04-20 · Anna P. Meyer, Aws Albarghouthi, Loris D'Antoni

We introduce dataset multiplicity, a way to study how inaccuracies, uncertainty, and social bias in training datasets impact test-time predictions. The dataset multiplicity framework asks a counterfactual question of what the set of resultant models (and associated test-time predictions) would be if we could somehow access all hypothetical, unbiased versions of the dataset. We discuss how to use this framework to encapsulate various sources of uncertainty in datasets' factualness, including systemic social bias, data collection practices, and noisy labels or features. We show how to exactly analyze the impacts of dataset multiplicity for a specific model architecture and type of uncertainty: linear models with label errors. Our empirical analysis shows that real-world datasets, under reasonable assumptions, contain many test samples whose predictions are affected by dataset multiplicity. Furthermore, the choice of domain-specific dataset multiplicity definition determines what samples are affected, and whether different demographic groups are disparately impacted. Finally, we discuss implications of dataset multiplicity for machine learning practice and research, including considerations for when model outcomes should not be trusted.

📄 PDF Abstract BibTeX arXiv:2304.10655

Code (1)

annapmeyer/linear-bias-certification 공식 구현 pytorch

Tasks

counterfactual

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

2025-05-26 · Sanghyuk Chun

Multimodal learning has seen remarkable progress, particularly with the emergence of large-scale pre-training across various modalities. However, most current approaches are built on the assumption of a deterministic, on…

Position

Predictive Multiplicity in Classification

2019-09-14 · ICML 2020 1 · Charles T. Marx, Flavio du Pin Calmon, Berk Ustun

Prediction problems often admit competing models that perform almost equally well. This effect challenges key assumptions in machine learning when competing models assign conflicting predictions. In this paper, we define…

ClassificationGeneral ClassificationModel SelectionPrediction

Accounting for multiplicity in machine learning benchmark performance

2023-03-10 · Kajsa Møllersen, Einar Holsbø

Machine learning methods are commonly evaluated and compared by their performance on data sets from public repositories. This allows for multiple methods, oftentimes several thousands, to be evaluated under identical con…

Multi-Target Multiplicity: Flexibility and Fairness in Target Specification under Resource Constraints

2023-06-23 · Jamelle Watson-Daniels, Solon Barocas, Jake M. Hofman, Alexandra Chouldechova

Prediction models have been widely adopted as the basis for decision-making in domains as diverse as employment, education, lending, and health. Yet, few real world problems readily present themselves as precisely formul…

Decision MakingFairness

Data as a Lever: A Neighbouring Datasets Perspective on Predictive Multiplicity

2025-10-24 · Prakhar Ganesh, Hsiang Hsu, Golnoosh Farnadi arxiv

Multiplicity, the existence of equally good yet competing models, has received growing attention in recent years. While prior work has emphasized modelling choices, the critical role of data in shaping multiplicity has b…

Active Learning