Improving OOD Generalization with Causal Invariant Transformations
In real-world applications, it is important and desirable to learn a model that performs well on out-of-distribution (OOD) data. Recently, causality has become a powerful tool to tackle the OOD generalization problem, with the core idea resting on the causal mechanism that is invariant across the domains of interest. To leverage the generally unknown causal mechanism, existing works assume the linear form of causal feature or require sufficiently many and diverse training domains, which are usually restrictive in practice. In this work, we obviate these assumptions and tackle the OOD problem without explicitly recovering the causal feature. Our approach is based on transformations that modify the non-causal feature but leave the causal part unchanged, which can be either obtained from prior knowledge or learned from the training data. Under the setting of invariant causal mechanism, we theoretically show that if all such transformations are available, then we can learn a minimax optimal model across the domains using only single domain data. Noticing that knowing a complete set of these causal invariant transformations may be impractical, we further show that it suffices to know only an appropriate subset of these transformations. Based on the theoretical findings, a regularized training procedure is proposed to improve the OOD generalization capability. Extensive experimental results on both synthetic and real datasets verify the effectiveness of the proposed algorithm, even with only a few causal invariant transformations.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Out-of-distribution Generalization with Causal Invariant Transformations
In real-world applications, it is important and desirable to learn a model that performs well on out-of-distribution (OOD) data. Recently, causality has become a powerful tool to tackle the OOD generalization problem, wi…
Out-of-Distribution GeneralizationCausality-inspired Latent Feature Augmentation for Single Domain Generalization
Single domain generalization (Single-DG) intends to develop a generalizable model with only one single training domain to perform well on other unknown target domains. Under the domain-hungry configuration, how to expand…
Domain GeneralizationGeneralization Error of Invariant Classifiers
This paper studies the generalization error of invariant classifiers. In particular, we consider the common scenario where the classification task is invariant to certain transformations of the input, and that the classi…
Nonlinear Invariant Risk Minimization: A Causal Approach
Due to spurious correlations, machine learning systems often fail to generalize to environments whose distributions differ from the ones used at training time. Prior work addressing this, either explicitly or implicitly,…
BIG-bench Machine LearningRepresentation LearningInvariant Causal Representation Learning for Out-of-Distribution Generalization
Due to spurious correlations, machine learning systems often fail to generalize to environments whose distributions differ from the ones used at training time. Prior work addressing this, either explicitly or implicitly,…
Out-of-Distribution GeneralizationRepresentation Learning