paper-with-me

홈 › Papers

CCSD: Cross-Modal Compositional Self-Distillation for Robust Brain Tumor Segmentation with Missing Modalities

2025-11-18 · Dongqing Xie, Yonghuang Wu, Zisheng Ai, Jun Min, Zhencun Jiang, Shaojin Geng, Lei Wang arxiv

The accurate segmentation of brain tumors from multi-modal MRI is critical for clinical diagnosis and treatment planning. While integrating complementary information from various MRI sequences is a common practice, the frequent absence of one or more modalities in real-world clinical settings poses a significant challenge, severely compromising the performance and generalizability of deep learning-based segmentation models. To address this challenge, we propose a novel Cross-Modal Compositional Self-Distillation (CCSD) framework that can flexibly handle arbitrary combinations of input modalities. CCSD adopts a shared-specific encoder-decoder architecture and incorporates two self-distillation strategies: (i) a hierarchical modality self-distillation mechanism that transfers knowledge across modality hierarchies to reduce semantic discrepancies, and (ii) a progressive modality combination distillation approach that enhances robustness to missing modalities by simulating gradual modality dropout during training. Extensive experiments on public brain tumor segmentation benchmarks demonstrate that CCSD achieves state-of-the-art performance across various missing-modality scenarios, with strong generalization and stability.

📄 PDF Abstract BibTeX arXiv:2511.14599

Code (0)

등록된 구현이 없습니다.

Tasks

Brain Tumor Segmentation

Similar Papers 제목 키워드 기반

Distilling Audio-Visual Knowledge by Compositional Contrastive Learning

2021-04-22 · CVPR 2021 1 · Yanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan 외

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous m…

Audio Taggingaudio-visual learningContrastive LearningKnowledge Distillation+3

Semantic Composition in Visually Grounded Language Models

2023-05-15 · Rohan Pandey

What is sentence meaning and its ideal representation? Much of the expressive power of human language derives from semantic composition, the mind's ability to represent meaning hierarchically & relationally over constitu…

Image CaptioningInductive BiasPhilosophyQuestion Answering+6

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

2026-08-26 · Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki 외 arxiv

Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this issue by augmenting …

Text Retrieval

Students taught by multimodal teachers are superior action recognizers

2022-10-09 · Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens 외

The focal point of egocentric video understanding is modelling hand-object interactions. Standard models -- CNNs, Vision Transformers, etc. -- which receive RGB frames as input perform well, however, their performance im…

Action RecognitionKnowledge DistillationObjectOptical Flow Estimation+1

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

2024-12-02 · CVPR 2025 1 · Sanghwan Kim, Rui Xiao, Mariana-Iuliana Georgescu, Stephan Alaniz 외

Vision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly o…

Self-Supervised LearningSemantic SegmentationUnsupervised Semantic Segmentation with Language-image Pre-trainingZero-Shot Cross-Modal Retrieval+1