paper-with-me

홈 › Papers

Understanding and Overcoming Cross-modal Fusion Bias in Multimodal Anomaly Detection From A Fisher Information Perspective

2026-08-02 · Kaifang Long, Lianbo Ma, Liming Liu, Guoyang Xie arxiv

Current advancements in Multimodal Anomaly Detection (MAD) are largely driven by enhancing multimodal fusion, particularly through the integration of RGB and Depth data for richer anomaly representation. However, less attention was devoted to analyzing the role of cross-modal fusion bias, a well-known challenge in multimodal learning, in MAD. This gap motivates a key question: can we overcome this bias to break the performance bottleneck of current work? In this paper, we first analyze the impact of cross-modal fusion bias in MAD via the Fisher Information Matrix. Then, grounded in these findings, we propose UCFB, a simple yet effective plug-and-play framework designed to mitigate cross-modal fusion bias in MAD. It achieves this by jointly employing Fisher-information-guided dynamic calibration to adjust modality-specific regularization weights and canonical similarity analysis to improve inter-modal interactions. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets demonstrate that UCFB achieves consistent improvements in single-class, multi-class, and few-shot settings.

📄 PDF Abstract BibTeX arXiv:2608.00986

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly Detection

Similar Papers 제목 키워드 기반

Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective

2024-03-27 · Meiqi Chen, Yixin Cao, Yan Zhang, Chaochao Lu

Recent advancements in Large Language Models (LLMs) have facilitated the development of Multimodal LLMs (MLLMs). Despite their impressive capabilities, MLLMs often suffer from over-reliance on unimodal biases (e.g., lang…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

MASS: Overcoming Language Bias in Image-Text Matching

2025-01-20 · Jiwan Chung, Seungwon Lim, Sangkyu Lee, Youngjae Yu

Pretrained visual-language models have made significant advancements in multimodal tasks, including image-text retrieval. However, a major challenge in image-text matching lies in language bias, where models predominantl…

Image-text matchingImage-text RetrievalMultimodal AssociationRetrieval+2

Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization

2025-11-25 · Xiaohan Wang, Zhangtao Cheng, Ting Zhong, Leiting Chen 외 arxiv

Weight Averaging (WA) has emerged as a powerful technique for enhancing generalization by promoting convergence to a flat loss landscape, which correlates with stronger out-of-distribution performance. However, applying …

Domain Generalization

Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model

2025-08-31 · Yifei She, Huangxuan Wu arxiv

Multimodal Large Language Models (MLLMs) have made significant progress in bridging visual perception with high-level textual reasoning. However, they face a fundamental contradiction: while excelling at complex semantic…

M$^3$amba: CLIP-driven Mamba Model for Multi-modal Remote Sensing Classification

2025-03-09 · Mingxiang Cao, Weiying Xie, Xin Zhang, Jiaqing Zhang 외

Multi-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fusion methods in the field of remote sens…

Computational EfficiencyHyperspectral Image Classificationimage-classificationImage Classification+1