Predictive Heterogeneity: Measures and Applications
As an intrinsic and fundamental property of big data, data heterogeneity exists in a variety of real-world applications, such as precision medicine, autonomous driving, financial applications, etc. For machine learning algorithms, the ignorance of data heterogeneity will greatly hurt the generalization performance and the algorithmic fairness, since the prediction mechanisms among different sub-populations are likely to differ from each other. In this work, we focus on the data heterogeneity that affects the prediction of machine learning models, and firstly propose the \emph{usable predictive heterogeneity}, which takes into account the model capacity and computational constraints. We prove that it can be reliably estimated from finite data with probably approximately correct (PAC) bounds. Additionally, we design a bi-level optimization algorithm to explore the usable predictive heterogeneity from data. Empirically, the explored heterogeneity provides insights for sub-population divisions in income prediction, crop yield prediction and image classification tasks, and leveraging such heterogeneity benefits the out-of-distribution generalization performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous DrivingCrop Yield PredictionFairnessimage-classificationImage ClassificationOut-of-Distribution GeneralizationPredictionSimilar Papers 제목 키워드 기반
On Information-Theoretic Measures of Predictive Uncertainty
Reliable estimation of predictive uncertainty is crucial for machine learning applications, particularly in high-stakes scenarios where hedging against risks is essential. Despite its significance, a consensus on the cor…
Out-of-Distribution DetectionSemiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity
We develop semiparametrically efficient inference for kernel measures of noise heterogeneity in additive noise models. In many applications, the regression function is estimated using flexible machine learning methods. D…
Forecasting with panel data: Estimation uncertainty versus parameter heterogeneity
We provide a comprehensive examination of the predictive performance of panel forecasting methods based on individual, pooling, fixed effects, and empirical Bayes estimation, and propose optimal weights for forecast comb…
Heterogeneous Image-based Classification Using Distributional Data Analysis
Diagnostic imaging has gained prominence as potential biomarkers for early detection and diagnosis in a diverse array of disorders including cancer. However, existing methods routinely face challenges arising from variou…
ClassificationDiagnosticquantile regressionSpecificityProperties of fairness measures in the context of varying class imbalance and protected group ratios
Society is increasingly relying on predictive models in fields like criminal justice, credit risk management, or hiring. To prevent such automated systems from discriminating against people belonging to certain groups, f…
Fairness