paper-with-me

Papers

Unbiased Dynamic Multimodal Fusion

2026-03-20 · Shicai Wei, Kaijie Zhang, Luyi Chen, Tao He, Guiduo Duan arxiv

Traditional multimodal methods often assume static modality quality, which limits their adaptability in dynamic real-world scenarios. Thus, dynamical multimodal methods are proposed to assess modality quality and adjust their contribution accordingly. However, they typically rely on empirical metrics, failing to measure the modality quality when noise levels are extremely low or high. Moreover, existing methods usually assume that the initial contribution of each modality is the same, neglecting the intrinsic modality dependency bias. As a result, the modality hard to learn would be doubly penalized, and the performance of dynamical fusion could be inferior to that of static fusion. To address these challenges, we propose the Unbiased Dynamic Multimodal Learning (UDML) framework. Specifically, we introduce a noise-aware uncertainty estimator that adds controlled noise to the modality data and predicts its intensity from the modality feature. This forces the model to learn a clear correspondence between feature corruption and noise level, allowing accurate uncertainty measure across both low- and high-noise conditions. Furthermore, we quantify the inherent modality reliance bias within multimodal networks via modality dropout and incorporate it into the weighting mechanism. This eliminates the dual suppression effect on the hard-to-learn modality. Extensive experiments across diverse multimodal benchmark tasks validate the effectiveness, versatility, and generalizability of the proposed UDML. The code is available at https://github.com/shicaiwei123/UDML.

📄 PDF Abstract BibTeX arXiv:2603.19681

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation

2024-05-24 · Bingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao 외

Multimodal semantic segmentation is a pivotal component of computer vision and typically surpasses unimodal methods by utilizing rich information set from various sources.Current models frequently adopt modality-specific…

SegmentationSemantic Segmentation

Doubly Stochastic Models: Learning with Unbiased Label Noises and Inference Stability

2023-04-01 · Haoyi Xiong, Xuhong LI, Boyang Yu, Zhanxing Zhu 외

Random label noises (or observational noises) widely exist in practical machine learning settings. While previous studies primarily focus on the affects of label noises to the performance of learning, our work intends to…

Provable Dynamic Fusion for Low-Quality Multimodal Data

2023-06-03 · Qingyang Zhang, Haitao Wu, Changqing Zhang, QinGhua Hu 외

The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-…

Predictive Dynamic Fusion

2024-06-07 · Bing Cao, Yinan Xia, Yi Ding, Changqing Zhang 외

Multimodal fusion is crucial in joint decision-making systems for rendering holistic judgments. Since multimodal data changes in open environments, dynamic fusion has emerged and achieved remarkable progress in numerous …

Decision Making

VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization

2020-03-12 · Kaichen Zhou, Changhao Chen, Bing Wang, Muhamad Risqi U. Saputra 외

Recent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missin…

Camera LocalizationCamera RelocalizationVisual Localization