paper-with-me

홈 › Papers

Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization

2025-11-25 · Xiaohan Wang, Zhangtao Cheng, Ting Zhong, Leiting Chen, Fan Zhou arxiv

Weight Averaging (WA) has emerged as a powerful technique for enhancing generalization by promoting convergence to a flat loss landscape, which correlates with stronger out-of-distribution performance. However, applying WA directly to multi-modal domain generalization (MMDG) is challenging: differences in optimization speed across modalities lead WA to overfit to faster-converging ones in early stages, suppressing the contribution of slower yet complementary modalities, thereby hindering effective modality fusion and skewing the loss surface toward sharper, less generalizable minima. To address this issue, we propose MBCD, a unified collaborative distillation framework that retains WA's flatness-inducing advantages while overcoming its shortcomings in multi-modal contexts. MBCD begins with adaptive modality dropout in the student model to curb early-stage bias toward dominant modalities. A gradient consistency constraint then aligns learning signals between uni-modal branches and the fused representation, encouraging coordinated and smoother optimization. Finally, a WA-based teacher conducts cross-modal distillation by transferring fused knowledge to each uni-modal branch, which strengthens cross-modal interactions and steer convergence toward flatter solutions. Extensive experiments on MMDG benchmarks show that MBCD consistently outperforms existing methods, achieving superior accuracy and robustness across diverse unseen domains.

📄 PDF Abstract BibTeX arXiv:2511.20258

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization

Similar Papers 제목 키워드 기반

Modality-Balanced Learning for Multimedia Recommendation

2024-07-26 · Jinghao Zhang, Guofan Liu, Qiang Liu, Shu Wu 외

Many recommender models have been proposed to investigate how to incorporate multimodal content information into traditional collaborative filtering framework effectively. The use of multimodal information is expected to…

Collaborative FilteringcounterfactualCounterfactual InferenceKnowledge Distillation+2

PASSION: Towards Effective Incomplete Multi-Modal Medical Image Segmentation with Imbalanced Missing Rates

2024-07-20 · Junjie Shi, Caozhi Shang, Zhaobin Sun, Li Yu 외

Incomplete multi-modal image segmentation is a fundamental task in medical imaging to refine deployment efficiency when only partial modalities are available. However, the common practice that complete-modality data is v…

Image SegmentationMedical Image SegmentationMulti-modal image segmentationSemantic Segmentation

Overcoming Uncertain Incompleteness for Robust Multimodal Sequential Diagnosis Prediction via Curriculum Data Erasing Guided Knowledge Distillation

2024-07-28 · Heejoon Koo

In this paper, we present NECHO v2, a novel framework designed to enhance the predictive accuracy of multimodal sequential patient diagnoses under uncertain missing visit sequences, a common challenge in real clinical se…

Knowledge DistillationSequential Diagnosis

CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification

2025-11-19 · Zhenyu Cui, Jiahuan Zhou, Yuxin Peng arxiv

Lifelong person Re-IDentification (LReID) aims to match the same person employing continuously collected individual data from different scenarios. To achieve continuous all-day person matching across day and night, Visib…

Person Re-IdentificationKnowledge Distillation

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation

2024-12-14 · Yang Yang, Wenjuan Xi, Luping Zhou, Jinhui Tang

Vision-language retrieval aims to search for similar instances in one modality based on queries from another modality. The primary objective is to learn cross-modal matching representations in a latent common space. Actu…

Cross-Modal RetrievalRetrieval