paper-with-me

홈 › Papers

Learning to Rebalance Multi-Modal Optimization by Adaptively Masking Subnetworks

2024-04-12 · Yang Yang, Hongpeng Pan, Qing-Yuan Jiang, Yi Xu, Jinghui Tang

Multi-modal learning aims to enhance performance by unifying models from various modalities but often faces the "modality imbalance" problem in real data, leading to a bias towards dominant modalities and neglecting others, thereby limiting its overall effectiveness. To address this challenge, the core idea is to balance the optimization of each modality to achieve a joint optimum. Existing approaches often employ a modal-level control mechanism for adjusting the update of each modal parameter. However, such a global-wise updating mechanism ignores the different importance of each parameter. Inspired by subnetwork optimization, we explore a uniform sampling-based optimization strategy and find it more effective than global-wise updating. According to the findings, we further propose a novel importance sampling-based, element-wise joint optimization method, called Adaptively Mask Subnetworks Considering Modal Significance(AMSS). Specifically, we incorporate mutual information rates to determine the modal significance and employ non-uniform adaptive sampling to select foreground subnetworks from each modality for parameter updates, thereby rebalancing multi-modal learning. Additionally, we demonstrate the reliability of the AMSS strategy through convergence analysis. Building upon theoretical insights, we further enhance the multi-modal mask subnetwork strategy using unbiased estimation, referred to as AMSS+. Extensive experiments reveal the superiority of our approach over comparison methods.

📄 PDF Abstract BibTeX arXiv:2404.08347

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-Spoofing

2024-02-29 · CVPR 2024 1 · Xun Lin, Shuai Wang, Rizhao Cai, Yizhong Liu 외

Face Anti-Spoofing (FAS) is crucial for securing face recognition systems against presentation attacks. With advancements in sensor manufacture and multi-modal learning techniques, many multi-modal FAS approaches have em…

Domain GeneralizationFace Anti-SpoofingFace Recognition

When Fusion Fails: Corruption-Aware Rebalanced Fusion for Multi-Modal Medical Image Segmentation

2026-09-09 · Yuchen Pei, Xiaoyu Hu, Yixiong Zou, Dingwen Hu 외 arxiv

Multi-modal medical image segmentation leverages complementary diagnostic information, yet fusion can underperform single-modality baselines when spatially aligned inputs differ in quality. Here, "corruption" primarily d…

Medical Image Segmentation

Rebalanced Multimodal Learning with Data-aware Unimodal Sampling

2025-03-05 · QingYuan Jiang, Zhouyang Chi, Xiao Ma, Qirong Mao 외

To address the modality learning degeneration caused by modality imbalance, existing multimodal learning~(MML) approaches primarily attempt to balance the optimization process of each modality from the perspective of mod…

Reinforcement Learning (RL)

Re-M3Dr: Rebalanced MultiModal Mean Deviation Regression

2026-05-26 · Haojie Yin, Chengcheng Feng, Tianyi Liu, Tianqi Zhang 외 arxiv

Mean Deviation (MD) is a critical metric for assessing visual field loss in ophthalmology. While previous work has focused solely on predicting MD from Optical Coherence Tomography (OCT), it is intuitive to assume that c…

Contrastive Learning

Reallocating Attention Across Layers to Reduce Multimodal Hallucination

2025-10-11 · Haolang Lu, Bolun Chu, WeiYe Fu, Guoshun Nan 외 arxiv

Multimodal large reasoning models (MLRMs) often suffer from hallucinations that stem not only from insufficient visual grounding but also from imbalanced allocation between perception and reasoning processes. Building up…

Multimodal ReasoningVisual Grounding