Classification under Nuisance Parameters and Generalized Label Shift in Likelihood-Free Inference
An open scientific challenge is how to classify events with reliable measures of uncertainty, when we have a mechanistic model of the data-generating process but the distribution over both labels and latent nuisance parameters is different between train and target data. We refer to this type of distributional shift as generalized label shift (GLS). Direct classification using observed data $\mathbf{X}$ as covariates leads to biased predictions and invalid uncertainty estimates of labels $Y$. We overcome these biases by proposing a new method for robust uncertainty quantification that casts classification as a hypothesis testing problem under nuisance parameters. The key idea is to estimate the classifier's receiver operating characteristic (ROC) across the entire nuisance parameter space, which allows us to devise cutoffs that are invariant under GLS. Our method effectively endows a pre-trained classifier with domain adaptation capabilities and returns valid prediction sets while maintaining high power. We demonstrate its performance on two challenging scientific problems in biology and astroparticle physics with data from realistic mechanistic models.
Code (1)
Tasks
Domain AdaptationUncertainty QuantificationvalidSimilar Papers 제목 키워드 기반
Deeper Understanding of Black-box Predictions via Generalized Influence Functions
Influence functions (IFs) elucidate how training data changes model behavior. However, the increasing size and non-convexity in large-scale models make IFs inaccurate. We suspect that the fragility comes from the first-o…
Influence ApproximationPhilosophyOut-of-distribution Generalization in the Presence of Nuisance-Induced Spurious Correlations
In many prediction problems, spurious correlations are induced by a changing relationship between the label and a nuisance variable that is also correlated with the covariates. For example, in classifying animals in natu…
Out-of-Distribution GeneralizationX-ray ClassificationOrthogonal Random Forest for Causal Inference
We propose the orthogonal random forest, an algorithm that combines Neyman-orthogonality to reduce sensitivity with respect to estimation error of nuisance parameters with generalized random forests (Athey et al., 2017)-…
Causal InferenceSemi-Supervised Sparse Representation Based Classification for Face Recognition with Insufficient Labeled Samples
This paper addresses the problem of face recognition when there is only few, or even only a single, labeled examples of the face that we wish to recognize. Moreover, these examples are typically corrupted by nuisance var…
Face RecognitionGeneral ClassificationSparse Representation-based ClassificationTransfer Learning for Causal Effect Estimation
We present a Transfer Causal Learning (TCL) framework when target and source domains share the same covariate/feature spaces, aiming to improve causal effect estimation accuracy in limited data. Limited data is very comm…
regressionTransfer Learning