paper-with-me

Papers

Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)

2022-03-23 · Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang, Longbo Huang

Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network outperforms the jointly trained multi-modal network, which is counter-intuitive since multiple signals generally bring more information. This work provides a theoretical explanation for the emergence of such performance gap in neural networks for the prevalent joint training framework. Based on a simplified data distribution that captures the realistic property of multi-modal data, we prove that for the multi-modal late-fusion network with (smoothed) ReLU activation trained jointly by gradient descent, different modalities will compete with each other. The encoder networks will learn only a subset of modalities. We refer to this phenomenon as modality competition. The losing modalities, which fail to be discovered, are the origins where the sub-optimality of joint training comes from. Experimentally, we illustrate that modality competition matches the intrinsic behavior of late-fusion joint training.

📄 PDF Abstract BibTeX arXiv:2203.12221

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Boosting Multi-modal Model Performance with Adaptive Gradient Modulation

2023-08-15 · ICCV 2023 1 · Hong Li, Xingyu Li, Pengbo Hu, Yinuo Lei 외

While the field of multi-modal learning keeps growing fast, the deficiency of the standard joint training paradigm has become clear through recent studies. They attribute the sub-optimal performance of the jointly traine…

Attribute

Shaping Initial State Prevents Modality Competition in Multi-modal Fusion: A Two-stage Scheduling Framework via Fast Partial Information Decomposition

2025-09-25 · Jiaqi Tang, Yinsong Xu, Yang Liu, Qingchao Chen arxiv

Multi-modal fusion often suffers from modality competition during joint training, where one modality dominates the learning process, leaving others under-optimized. Overlooking the critical impact of the model's initial …

Boosting Multimodal Federated Learning via Chained Modality Optimization

2026-06-01 · Zixin Zhang, Fan Qi, Shuai Li, Xiaoshan Yang 외 arxiv

Multimodal Federated Learning (MMFL) enables privacy-preserving collaborative learning across decentralized clients with heterogeneous data and modality availability. However, most existing MMFL methods cast multimodal t…

Federated Learning

Detached and Interactive Multimodal Learning

2024-07-28 · Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junhong Liu 외

Recently, Multimodal Learning (MML) has gained significant interest as it compensates for single-modality limitations through comprehensive complementary information within multimodal data. However, traditional MML metho…

Transfer Learning

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

2026-01-06 · Yusheng Dai, Zehua Chen, Yuxuan Jiang, Baolong Gao 외 arxiv

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges…

Audio Generation