paper-with-me

홈 › Papers

M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model

2026-04-10 · Yihang Liu, Longzhen Yang, Jiaxiong Yang, Ying Wen, Lianghua He, Heng Tao Shen arxiv

Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse downstream clinical tasks. However, most existing MFMs suffer from information ambiguity that blends multimodal representations in a single embedding space, leading to the degradation of modality specificity and diversity. In this paper, we propose M-IDoL, a self-supervised MFM that introduces Information Decomposition for multimodal representation Learning via two objectives: i) maximizing inter-modality entropy by dispersing multimodal representations into separable Mixture-of-Experts (MoE) subspaces to achieve representation specificity across modalities; and ii) minimizing intra-modality uncertainty by performing fine-grained semantic discrimination within each MoE subspace to enrich representation diversity per modality. By pre-training on 1.15 million medical images, M-IDoL i) delivers superior generalization across 21 downstream clinical tasks, outperforming 20 foundation models on five imaging modalities (e.g., X-ray, fundus, OCT, dermoscopy and pathology), and ii) learns modality-specific and diverse representations, showing clearer separation of feature clusters across modalities and finer-grained feature discrimination within each modality.

📄 PDF Abstract BibTeX arXiv:2604.08936

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Intentional Deep Overfit Learning (IDOL): A Novel Deep Learning Strategy for Adaptive Radiation Therapy

2021-04-23 · Jaehee Chun, Justin C. Park, Sven Olberg, You Zhang 외

In this study, we propose a tailored DL framework for patient-specific performance that leverages the behavior of a model intentionally overfitted to a patient-specific training dataset augmented from the prior informati…

Super-Resolution

Pareidolia Face Reenactment

2021-06-19 · CVPR 2021 1 · Linsen Song, Wayne Wu, Chaoyou Fu, Chen Qian 외

We present a new application direction named Pareidolia Face Reenactment, which is defined as animating a static illusory face to move in tandem with a human face in the video. For the large differences between parei…

Face ReenactmentTexture Synthesis

Everything's Talkin': Pareidolia Face Reenactment

2021-04-07 · Linsen Song, Wayne Wu, Chaoyou Fu, Chen Qian 외

We present a new application direction named Pareidolia Face Reenactment, which is defined as animating a static illusory face to move in tandem with a human face in the video. For the large differences between pareidoli…

Face ReenactmentTexture Synthesis

IDOL: Meeting Diverse Distribution Shifts with Prior Physics for Tropical Cyclone Multi-Task Estimation

2025-11-13 · Hanting Yan, Pan Mu, Shiqi Zhang, Yuchao Zhu 외 arxiv

Tropical Cyclone (TC) estimation aims to accurately estimate various TC attributes in real time. However, distribution shifts arising from the complex and dynamic nature of TC environmental fields, such as varying geogra…

IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation

2024-07-15 · Yuanhao Zhai, Kevin Lin, Linjie Li, Chung-Ching Lin 외

Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation methods may not generalize well to synth…

DenoisingDepth EstimationMonocular Depth EstimationVideo Denoising+1