When Sample Selection Bias Precipitates Model Collapse
The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes distributional tails and homogenizes outputs. Data selection is widely viewed as a remedy, yet its reliability depends critically on the reference distribution used by the verifier. We show that in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection itself becomes biased. This situation naturally arises in low-resource data silos such as healthcare consortia or proprietary financial institutions, where raw data cannot be pooled and local references are inherently incomplete. As a result, selection preferentially retains samples aligned with the local manifold while pruning globally relevant tail modes, turning from a safeguard against collapse into a mechanism that precipitates it. We theoretically prove that such siloed selection accelerates collapse and induces power-law diversity decay. As an initial mitigation, we construct Wasserstein proxy references from multiple silos without sharing raw data. Empirical results confirm that local-reference selection fails on skewed distributions, whereas collaborative proxy references mitigate diversity degradation, suggesting that recursive synthetic-data pipelines require particular caution when real-data coverage is fragmented or scarce.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Fair Classifier Embracing Triplet Collapse
In this paper, we study the behaviour of the triplet loss and show that it can be exploited to limit the biases created and perpetuated by machine learning models. Our fair classifier uses the collapse of the triplet los…
TripletInstrumental Variables with Treatment-Induced Selection: Exact Bias Results
Instrumental variables (IV) estimation suffers selection bias when the analysis conditions on the treatment. Judea Pearl's early graphical definition of instrumental variables explicitly prohibited conditioning on the tr…
Selection biasA Probabilistic Perspective on Model Collapse
In recent years, model collapse has become a critical issue in language model training, making it essential to understand the underlying mechanisms driving this phenomenon. In this paper, we investigate recursive paramet…
modelEvaluating the Fairness of Neural Collapse in Medical Image Classification
Deep learning has achieved impressive performance across various medical imaging tasks. However, its inherent bias against specific groups hinders its clinical applicability in equitable healthcare systems. A recently di…
Deep LearningFairnessimage-classificationImage Classification+1On Prediction Feature Assignment in the Heckman Selection Model
Under missing-not-at-random (MNAR) sample selection bias, the performance of a prediction model is often degraded. This paper focuses on one classic instance of MNAR sample selection bias where a subset of samples have n…
PredictionSelection bias