paper-with-me

Papers

Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

2025-08-07 · Xusheng Liang, Lihua Zhou, Nianxin Li, Miao Xu, Ziyang Song, Dong Yi, Jinlin Wu, Jiawei Ma, Hongbin Liu, Zhen Lei, Jiebo Luo arxiv

Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging remains challenging due to the high variability and complexity of medical data. Specifically, medical images often exhibit significant domain shifts caused by various confounders, including equipment differences, procedure artifacts, and imaging modes, which can lead to poor generalization when models are applied to unseen domains. To address this limitation, we propose Multimodal Causal-Driven Representation Learning (MCDRL), a novel framework that integrates causal inference with the VLM to tackle domain generalization in medical image segmentation. MCDRL is implemented in two steps: first, it leverages CLIP's cross-modal capabilities to identify candidate lesion regions and construct a confounder dictionary through text prompts, specifically designed to represent domain-specific variations; second, it trains a causal intervention network that utilizes this dictionary to identify and eliminate the influence of these domain-specific variations while preserving the anatomical structural information critical for segmentation tasks. Extensive experiments demonstrate that MCDRL consistently outperforms competing methods, yielding superior segmentation accuracy and exhibiting robust generalizability.

📄 PDF Abstract BibTeX arXiv:2508.05008

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Image SegmentationRepresentation LearningDomain GeneralizationCausal Inference

Similar Papers 제목 키워드 기반

A tutorial on discovering and quantifying the effect of latent causal sources of multimodal EHR data

2025-10-15 · Marco Barbero-Mota, Eric V. Strobl, John M. Still, William W. Stead 외 arxiv

We provide an accessible description of a peer-reviewed generalizable causal machine learning pipeline to (i) discover latent causal sources of large-scale electronic health records observations, and (ii) quantify the so…

Robust Multimodal Representation Learning in Healthcare

2026-01-29 · Xiaoguang Zhu, Linxiao Gong, Lianlong Sun, Yang Liu 외 arxiv

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systemati…

Representation Learning

CAMO: Causality-Guided Adversarial Multimodal Domain Generalization for Crisis Classification

2025-12-08 · Pingchuan Ma, Chengshuai Zhao, Bohan Jiang, Saketh Vishnubhatla 외 arxiv

Crisis classification in social media aims to extract actionable disaster-related information from multimodal posts, which is a crucial task for enhancing situational awareness and facilitating timely emergency responses…

Representation LearningDomain Generalization

Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities

2025-09-06 · Xiaoguang Zhu, Lianlong Sun, Yang Liu, Pengyi Jiang 외 arxiv

Medical multimodal representation learning aims to integrate heterogeneous clinical data into unified patient representations to support predictive modeling, which remains an essential yet challenging task in the medical…

Representation Learning

Causal Representation Learning from Multimodal Biomedical Observations

2024-11-10 · Yuewen Sun, Lingjing Kong, Guangyi Chen, Loka Li 외

Prevalent in biomedical applications (e.g., human phenotype research), multimodal datasets can provide valuable insights into the underlying physiological mechanisms. However, current machine learning (ML) models designe…

Representation Learning