Towards a Theoretical Framework of Out-of-Distribution Generalization
Generalization to out-of-distribution (OOD) data is one of the central problems in modern machine learning. Recently, there is a surge of attempts to propose algorithms that mainly build upon the idea of extracting invariant features. Although intuitively reasonable, theoretical understanding of what kind of invariance can guarantee OOD generalization is still limited, and generalization to arbitrary out-of-distribution is clearly impossible. In this work, we take the first step towards rigorous and quantitative definitions of 1) what is OOD; and 2) what does it mean by saying an OOD problem is learnable. We also introduce a new concept of expansion function, which characterizes to what extent the variance is amplified in the test domains over the training domains, and therefore give a quantitative meaning of invariant features. Based on these, we prove OOD generalization error bounds. It turns out that OOD generalization largely depends on the expansion function. As recently pointed out by Gulrajani and Lopez-Paz (2020), any OOD learning algorithm without a model selection module is incomplete. Our theory naturally induces a model selection criterion. Extensive experiments on benchmark OOD datasets demonstrate that our model selection criterion has a significant advantage over baselines.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain GeneralizationModel SelectionOut-of-Distribution GeneralizationSimilar Papers 제목 키워드 기반
Temporal Domain Generalization with Drift-Aware Dynamic Neural Networks
Temporal domain generalization is a promising yet extremely challenging area where the goal is to learn models under temporally changing data distributions and generalize to unseen data distributions following the trends…
Domain GeneralizationDynamic neural networksGraph GenerationDistributionally Robust Graph Out-of-Distribution Recommendation via Diffusion Model
The distributionally robust optimization (DRO)-based graph neural network methods improve recommendation systems' out-of-distribution (OOD) generalization by optimizing the model's worst-case performance. However, these …
Graph Neural NetworkRecommendation SystemsModify Training Directions in Function Space to Reduce Generalization Error
We propose theoretical analyses of a modified natural gradient descent method in the neural network function space based on the eigendecompositions of neural tangent kernel and Fisher information matrix. We firstly prese…
PAC-Chernoff Bounds: Understanding Generalization in the Interpolation Regime
This paper introduces a distribution-dependent PAC-Chernoff bound that exhibits perfect tightness for interpolators, even within over-parameterized model classes. This bound, which relies on basic principles of Large Dev…
Data AugmentationTheoretical Grounding of Out-Of-Distribution Detection With Reinforcement Learning Optimizer
Out-of-distribution (OOD) detection in dynamic open-world environments requires a model to continually adapt to evolving data distributions while generalizing to covariate-shifted inputs and rejecting semantic-shifted OO…
Out-of-Distribution DetectionReinforcement LearningDomain Generalization