paper-with-me

홈 › Papers

DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion

2026-01-09 · Yiming Sun, Zifan Ye, Qinghua Hu, Pengfei Zhu arxiv

Multi-modal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space model have achieved satisfied performance with high computational efficiency, they tend to either over-prioritize infrared intensity at the cost of visible details, or conversely, preserve visible structure while diminishing thermal target salience. To overcome these challenges, we propose DIFF-MF, a novel difference-driven channel-spatial state space model for multi-modal image fusion. Our approach leverages feature discrepancy maps between modalities to guide feature extraction, followed by a fusion process across both channel and spatial dimensions. In the channel dimension, a channel-exchange module enhances channel-wise interaction through cross-attention dual state space modeling, enabling adaptive feature reweighting. In the spatial dimension, a spatial-exchange module employs cross-modal state space scanning to achieve comprehensive spatial fusion. By efficiently capturing global dependencies while maintaining linear computational complexity, DIFF-MF effectively integrates complementary multi-modal features. Experimental results on the driving scenarios and low-altitude UAV datasets demonstrate that our method outperforms existing approaches in both visual quality and quantitative evaluation.

📄 PDF Abstract BibTeX arXiv:2601.05538

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Enhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learning

2020-03-09 · Rongzhi Gu, Shi-Xiong Zhang, Lian-Wu Chen, Yong Xu 외

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial fea…

Speech Separation

A spatial hue similarity measure for assessment of colourisation

2020-11-03 · Seán Mullery, Paul F. Whelan

Automatic colourisation of grey-scale images is an ill-posed multi-modal problem. Where full-reference images exist, objective performance measures rely on pixel-difference techniques such as MSE and PSNR. These measures…

SSIM

End-to-End Multi-Channel Speech Separation

2019-05-15 · Rongzhi Gu, Jian Wu, Shi-Xiong Zhang, Lian-Wu Chen 외

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end-to-end model for multi-channel speech s…

Speech Separation

A Remote Sensing Image Change Detection Method Integrating Layer Exchange and Channel-Spatial Differences

2025-01-19 · Sijun Dong, Fangcheng Zuo, Geng Chen, Siming Fu 외

Change detection in remote sensing imagery is a critical technique for Earth observation, primarily focusing on pixel-level segmentation of change regions between bi-temporal images. The essence of pixel-level change det…

Change DetectionEarth Observation

Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification

2024-07-02 · Shulei Qiu, Wanqi Yang, Ming Yang

Our research focuses on few-shot fine-grained image classification, which faces two major challenges: appearance similarity of fine-grained objects and limited number of samples. To preserve the appearance details of ima…

Fine-Grained Image Classificationimage-classificationImage Classification