paper-with-me

홈 › Papers

Learning Invariant Causal Mechanism from Vision-Language Models

2024-05-24 · Zeen Song, Siyu Zhao, Xingyu Zhang, Jiangmeng Li, Changwen Zheng, Wenwen Qiang

Large-scale pre-trained vision-language models such as CLIP have been widely applied to a variety of downstream scenarios. In real-world applications, the CLIP model is often utilized in more diverse scenarios than those encountered during its training, a challenge known as the out-of-distribution (OOD) problem. However, our experiments reveal that CLIP performs unsatisfactorily in certain domains. Through a causal analysis, we find that CLIP's current prediction process cannot guarantee a low OOD risk. The lowest OOD risk can be achieved when the prediction process is based on invariant causal mechanisms, i.e., predicting solely based on invariant latent factors. However, theoretical analysis indicates that CLIP does not identify these invariant latent factors. Therefore, we propose the Invariant Causal Mechanism for CLIP (CLIP-ICM), a framework that first identifies invariant latent factors using interventional data and then performs invariant predictions across various domains. Our method is simple yet effective, without significant computational overhead. Experimental results demonstrate that CLIP-ICM significantly improves CLIP's performance in OOD scenarios.

📄 PDF Abstract BibTeX arXiv:2405.15289

Code (0)

등록된 구현이 없습니다.

Tasks

Causal InferenceDecision Making

Methods 이 논문이 사용한 방법론

Causal inference Causal inference is the process of drawing a conclusion about a causal connection based on the conditions of the occurrence of an effect. The main difference between causal…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Improving OOD Generalization with Causal Invariant Transformations

2021-09-29 · Ruoyu Wang, Mingyang Yi, Shengyu Zhu, Zhitang Chen

In real-world applications, it is important and desirable to learn a model that performs well on out-of-distribution (OOD) data. Recently, causality has become a powerful tool to tackle the OOD generalization problem, wi…

Out-of-distribution Generalization with Causal Invariant Transformations

2022-03-22 · CVPR 2022 1 · Ruoyu Wang, Mingyang Yi, Zhitang Chen, Shengyu Zhu

In real-world applications, it is important and desirable to learn a model that performs well on out-of-distribution (OOD) data. Recently, causality has become a powerful tool to tackle the OOD generalization problem, wi…

Out-of-Distribution Generalization

BiPrompt: Bilateral Prompt Optimization for Visual and Textual Debiasing in Vision-Language Models

2026-01-05 · Sunny Gupta, Shounak Das, Amit Sethi arxiv

Vision language foundation models such as CLIP exhibit impressive zero-shot generalization yet remain vulnerable to spurious correlations across visual and textual modalities. Existing debiasing approaches often address …

Zero-shot GeneralizationTest-time Adaptation

Can Large Language Models Learn Independent Causal Mechanisms?

2024-02-04 · Gaël Gendron, Bao Trung Nguyen, Alex Yuxuan Peng, Michael Witbrock 외

Despite impressive performance on language modelling and complex reasoning tasks, Large Language Models (LLMs) fall short on the same tasks in uncommon settings or with distribution shifts, exhibiting a lack of generalis…

Language Modelling

Invariant Structure Learning for Better Generalization and Causal Explainability

2022-06-13 · Yunhao Ge, Sercan Ö. Arik, Jinsung Yoon, Ao Xu 외

Learning the causal structure behind data is invaluable for improving generalization and obtaining high-quality explanations. We propose a novel framework, Invariant Structure Learning (ISL), that is designed to improve …

Self-Supervised Learning