paper-with-me

홈 › Papers

Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection

2025-11-19 · YiKang Shao, Tao Shi arxiv

Multimodal object detection has attracted significant attention in both academia and industry for its enhanced robustness. Although numerous studies have focused on improving modality fusion strategies, most neglect fusion degradation, and none provide a theoretical analysis of its underlying causes. To fill this gap, this paper presents a systematic theoretical investigation of fusion degradation in multimodal detection and identifies two key optimization deficiencies: (1) the gradients of unimodal branch backbones are severely suppressed under multimodal architectures, resulting in under-optimization of the unimodal branches; (2) disparities in modality quality cause weaker modalities to experience stronger gradient suppression, which in turn results in imbalanced modality learning. To address these issues, this paper proposes a Representation Space Constrained Learning with Modality Decoupling (RSC-MD) method, which consists of two modules. The RSC module and the MD module are designed to respectively amplify the suppressed gradients and eliminate inter-modality coupling interference as well as modality imbalance, thereby enabling the comprehensive optimization of each modality-specific backbone. Extensive experiments conducted on the FLIR, LLVIP, M3FD, and MFAD datasets demonstrate that the proposed method effectively alleviates fusion degradation and achieves state-of-the-art performance across multiple benchmarks. The code and training procedures will be released at https://github.com/yikangshao/RSC-MD.

📄 PDF Abstract BibTeX arXiv:2511.15433

Code (0)

등록된 구현이 없습니다.

Tasks

Object Detection

Similar Papers 제목 키워드 기반

Robust Multimodal Learning via Representation Decoupling

2024-07-05 · Shicai Wei, Yang Luo, Yuji Wang, Chunbo Luo

Multimodal learning robust to missing modality has attracted increasing attention due to its practicality. Existing methods tend to address it by learning a common subspace representation for different modality combinati…

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

2026-02-23 · Chunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu 외 arxiv

Multimodal Sentiment Analysis (MSA) integrates language, visual, and acoustic modalities to infer human sentiment. Most existing methods either focus on globally shared representations or modality-specific features, whil…

Multimodal Intent RecognitionMultimodal Sentiment Analysis

Decoupling Common and Unique Representations for Multimodal Self-supervised Learning

2023-09-11 · Yi Wang, Conrad M Albrecht, Nassim Ait Ali Braham, Chenying Liu 외

The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-mod…

Scene ClassificationSelf-Supervised LearningSemantic Segmentation

Anchored Alignment: Preventing Positional Collapse in Multimodal Recommender Systems

2026-03-13 · Yonghun Jeong, David Yoon Suk Kang, Yeon-Chang Lee arxiv

Multimodal recommender systems (MMRS) leverage images, text, and interaction signals to enrich item representations. However, recent alignment based MMRSs that enforce a unified embedding space often blur modality specif…

Multimodal RecommendationRepresentation Learning

Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs

2026-04-20 · Tingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang 외 arxiv

Although Knowledge Editing provides an efficient mechanism for updating the knowledge of Multimodal Large Language Models (MLLMs), we find that current paradigms still suffer from an important yet remain underexplored is…

knowledge editing