On the Need for a Language Describing Distribution Shifts: Illustrations on Tabular Datasets
Different distribution shifts require different algorithmic and operational
interventions. Methodological research must be grounded by the specific
shifts they address. Although nascent benchmarks provide a promising
empirical foundation, they \emph{implicitly} focus on covariate
shifts, and the validity of empirical findings depends on the type of shift,
e.g., previous observations on algorithmic performance can fail to be valid when
the $Y|X$ distribution changes. We conduct a thorough investigation of
natural shifts in 5 tabular datasets over 86,000 model configurations, and
find that $Y|X$-shifts are most prevalent. To encourage researchers to
develop a refined language for distribution shifts, we build
`WhyShift`, an empirical testbed of curated real-world shifts where
we characterize the type of shift we benchmark performance over. Since
$Y|X$-shifts are prevalent in tabular settings, we \emph{identify covariate
regions} that suffer the biggest $Y|X$-shifts and discuss implications for
algorithmic and data-based interventions. Our testbed highlights the
importance of future research that builds an understanding of why
distributions differ.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Are Layout-Infused Language Models Robust to Layout Distribution Shifts? A Case Study with Scientific Documents
Recent work has shown that infusing layout features into language models (LMs) improves processing of visually-rich documents such as scientific papers. Layout-infused LMs are often evaluated on documents with familiar l…
DiversityVerified Training for Counterfactual Explanation Robustness under Data Shift
Counterfactual explanations (CEs) enhance the interpretability of machine learning models by describing what changes to an input are necessary to change its prediction to a desired class. These explanations are commonly …
counterfactualCounterfactual ExplanationDr. FERMI: A Stochastic Distributionally Robust Fair Empirical Risk Minimization Framework
While training fair machine learning models has been studied extensively in recent years, most developed methods rely on the assumption that the training and test data have similar distributions. In the presence of distr…
FairnessClustering of illustrations by atmosphere using a combination of supervised and unsupervised learning
The distribution of illustrations on social media, such as Twitter and Pixiv has increased with the growing popularity of animation, games, and animated movies. The "atmosphere" of illustrations plays an important role i…
ClusteringDescribing Differences between Text Distributions with Natural Language
How do two distributions of texts differ? Humans are slow at answering this, since discovering patterns might require tediously reading through hundreds of samples. We propose to automatically summarize the differences b…
Binary ClassificationRe-Ranking