paper-with-me

홈 › Papers

Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning

2025-08-05 · Tianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi arxiv

Few-shot learning (FSL) often requires effective adaptation of models using limited labeled data. However, most existing FSL methods rely on entangled representations, requiring the model to implicitly recover the unmixing process to obtain disentangled representations using only limited supervision, which hinders effective adaptation. Recent theoretical studies show that multimodal contrastive learning methods, such as CLIP, can disentangle latent representations up to linear transformations. In light of this, we propose the Causal CLIP Adapter (CCA), a novel framework that explicitly disentangles visual features extracted from CLIP using unsupervised Independent Component Analysis (ICA). This removes the need to learn the unmixing process from the labeled data, thereby reducing the number of trainable parameters and mitigating overfitting. Taking a step further, while ICA can obtain visual disentangled representations, it may also disrupt CLIP's intra- and inter-modal alignment. To counteract this, CCA further leverages CLIP's inherent cross-modal alignment by enhancing it in two ways: unidirectionally, through fine-tuning a CLIP-based text classifier, and bidirectionally, via a cross-attention mechanism that enriches visual and textual representations through mutual interaction. Both unimodal and cross-modal classification outputs can be effectively combined linearly to improve classification accuracy. Extensive experiments on 11 benchmark datasets demonstrate that our method consistently outperforms state-of-the-art approaches in terms of few-shot performance and robustness to distributional shifts, while maintaining computational efficiency. Code will be available at https://github.com/tianjiao-j/CCA.

📄 PDF Abstract BibTeX arXiv:2508.03102

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyContrastive LearningFew-Shot Learning

Similar Papers 제목 키워드 기반

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

2026-06-08 · Zhou Du, Hamid Krim, Xiao Wu, Zhaoquan Yuan 외 arxiv

Recent advances in video multimodal models have significantly improved VideoQA performance. However, these systems often rely on spurious statistical correlations rather than answer-relevant causal evidence, resulting in…

CausalDisenSeg: A Causality-Guided Disentanglement Framework with Counterfactual Reasoning for Robust Brain Tumor Segmentation Under Missing Modalities

2026-04-15 · Bo Liu, Yulong Zou, Jin Hong arxiv

In clinical practice, the robustness of deep learning models for multimodal brain tumor segmentation is severely compromised by incomplete MRI data. This vulnerability stems primarily from modality bias, where models exp…

Brain Tumor Segmentation

Aligning Non-Causal Factors for Transformer-Based Source-Free Domain Adaptation

2023-11-27 · Sunandini Sanyal, Ashish Ramayee Asokan, Suvaansh Bhambri, Pradyumna YM 외

Conventional domain adaptation algorithms aim to achieve better generalization by aligning only the task-discriminative causal factors between a source and target domain. However, we find that retaining the spurious corr…

DisentanglementDomain AdaptationPrivacy PreservingSource-Free Domain Adaptation

Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models

2025-05-26 · Xinmiao Hu, Chun Wang, Ruihe An, ChenYu Shao 외

Multimodal Large Language Models (MLLMs) have demonstrated strong performance in visual understanding tasks, yet they often suffer from object hallucinations--generating descriptions of objects that are inconsistent with…

DisentanglementHallucinationLanguage ModelingLanguage Modelling+1

Disentangled Noisy Correspondence Learning

2024-08-10 · Zhuohang Dang, Minnan Luo, Jihong Wang, Chengyou Jia 외

Cross-modal retrieval is crucial in understanding latent correspondences across modalities. However, existing methods implicitly assume well-matched training data, which is impractical as real-world data inevitably invol…

cross-modal alignmentCross-Modal RetrievalDisentanglementMutual Information Estimation