paper-with-me

Papers

Balanced Multimodal Learning via Mutual Information

2025-11-02 · Rongrong Xie, Guido Sanguinetti arxiv

Multimodal learning has increasingly become a focal point in research, primarily due to its ability to integrate complementary information from diverse modalities. Nevertheless, modality imbalance, stemming from factors such as insufficient data acquisition and disparities in data quality, has often been inadequately addressed. This issue is particularly prominent in biological data analysis, where datasets are frequently limited, costly to acquire, and inherently heterogeneous in quality. Conventional multimodal methodologies typically fall short in concurrently harnessing intermodal synergies and effectively resolving modality conflicts. In this study, we propose a novel unified framework explicitly designed to address modality imbalance by utilizing mutual information to quantify interactions between modalities. Our approach adopts a balanced multimodal learning strategy comprising two key stages: cross-modal knowledge distillation (KD) and a multitask-like training paradigm. During the cross-modal KD pretraining phase, stronger modalities are leveraged to enhance the predictive capabilities of weaker modalities. Subsequently, our primary training phase employs a multitask-like learning mechanism, dynamically calibrating gradient contributions based on modality-specific performance metrics and intermodal mutual information. This approach effectively alleviates modality imbalance, thereby significantly improving overall multimodal model performance.

📄 PDF Abstract BibTeX arXiv:2511.00987

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

Neural Network Classifier as Mutual Information Evaluator

2021-06-19 · Zhenyue Qin, Dongwoo Kim, Tom Gedeon

Cross-entropy loss with softmax output is a standard choice to train neural network classifiers. We give a new view of neural network classifiers with softmax and cross-entropy as mutual information evaluators. We show t…

Form

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

2026-07-29 · Xuan Feng, Guihong Liu, Tianlong Gu, Shuai Zhao 외 arxiv

Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplified by imbalanced data and semantically inconsistent text-image pairs…

Balanced Diffusion-Guided Fusion for Multimodal Remote Sensing Classification

2025-09-27 · Hao Liu, Yongjie Zheng, Yuhan Kang, Mingyang Zhang 외 arxiv

Deep learning-based techniques for the analysis of multimodal remote sensing data have become popular due to their ability to effectively integrate complementary spatial, spectral, and structural information from differe…

Asymmetric Reinforcing against Multi-modal Representation Bias

2025-01-02 · Xiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang 외

The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the chall…

Robust Multimodal Semantic Segmentation with Balanced Modality Contributions

2025-09-29 · Jiaqi Tan, Xu Zheng, Fangyu Li, Yang Liu arxiv

Multimodal semantic segmentation enhances model robustness by exploiting cross-modal complementarities. However, existing methods often suffer from imbalanced modal dependencies, where overall performance degrades signif…

Semantic Segmentation