paper-with-me

홈 › Papers

Heckman-Corrected Epistemic Uncertainty: Selection on Unobservables Defeats Importance Weighting

2026-07-07 · Gunner Levi Howe arxiv

Training data for machine learning is routinely collected by a selection process the model never sees: loans are observed only when granted, outcomes only when a test was ordered. The standard fixes -- importance weighting, covariate-shift correction, MAR imputation -- assume selection is ignorable given observables. Econometrics solved the harder case in 1979: Heckman's two-equation model jointly fits a probit selection equation and an outcome equation linked through correlated errors, and the inverse-Mills-ratio term corrects for selection on unobservables, where importance weighting is structurally helpless. We instantiate this for deep epistemic uncertainty: a deep outcome network, a linear selection head, and a joint bivariate-normal likelihood over all units, ensembled for predictive variance. In a controlled generator where sampling probability depends on an unobservable correlated (rho up to 0.9) with the outcome noise, deep ensembles, MC dropout, and GP baselines are overconfident exactly where data was avoided: coverage of nominal-90% intervals falls to 64.4% at rho=0.9, and importance weighting with oracle propensities does not fix it (43.1%) -- reweighting corrects the covariate distribution, not the conditional bias E[y|x,selected] != E[y|x]. The Heckman correction restores coverage (88.9%) when the selection equation has an instrument -- a variable affecting selection but not the outcome -- and degrades measurably without one (40.3%); we chart this honesty curve rather than hide it. On real tabular data with induced MNAR selection, the corrected intervals are the best-calibrated (lowest region-ECE) non-oracle method in selected-against regions; baselines matching its raw coverage do so only by over-widening everywhere. Our estimators reproduce classic Stata output to seven digits. We state which identification regime a practitioner is in, and release the code.

📄 PDF Abstract BibTeX arXiv:2607.05806

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Prediction Feature Assignment in the Heckman Selection Model

2023-09-14 · Huy Mai, Xintao Wu

Under missing-not-at-random (MNAR) sample selection bias, the performance of a prediction model is often degraded. This paper focuses on one classic instance of MNAR sample selection bias where a subset of samples have n…

PredictionSelection bias

Heckman-Selection or Two-Part models for alcohol studies? Depends

2021-12-20 · Reka Sundaram-Stukel

Aims: To re-introduce the Heckman model as a valid empirical technique in alcohol studies. Design: To estimate the determinants of problem drinking using a Heckman and a two-part estimation model. Psychological and neuro…

Selection biasSurveyVocal Bursts Valence Prediction

Logit Disagreement: OoD Detection with Bayesian Neural Networks

2025-02-21 · Kevin Raina

Bayesian neural networks (BNNs), which estimate the full posterior distribution over model parameters, are well-known for their role in uncertainty quantification and its promising application in out-of-distribution dete…

Out-of-Distribution DetectionUncertainty QuantificationVariational Inference

Where to Measure: Epistemic Uncertainty-Based Sensor Placement with ConvCNPs

2025-11-27 · Feyza Eksen, Stefan Oehmcke, Stefan Lüdtke arxiv

Accurate sensor placement is critical for modeling spatio-temporal systems such as environmental and climate processes. Neural Processes (NPs), particularly Convolutional Conditional Neural Processes (ConvCNPs), provide …

Causal EpiNets: Precision-corrected Bounds on Individual Treatment Effects using Epistemic Neural Networks

2026-05-08 · Gandharv Patil, Keyi Tang, Raquel Aoki, Leo Guelman arxiv

Individual treatment effects are not point-identified from data. The Probability of Necessity and Sufficiency (PNS) circumvents this limitation by characterizing individual-level causality through intersection bounds der…