paper-with-me

홈 › Papers

CoMiX: Cross-Modal Fusion with Deformable Convolutions for HSI-X Semantic Segmentation

2024-11-13 · Xuming Zhang, Xingfa Gu, Qingjiu Tian, Lorenzo Bruzzone

Improving hyperspectral image (HSI) semantic segmentation by exploiting complementary information from a supplementary data type (referred to X-modality) is promising but challenging due to differences in imaging sensors, image content, and resolution. Current techniques struggle to enhance modality-specific and modality-shared information, as well as to capture dynamic interaction and fusion between different modalities. In response, this study proposes CoMiX, an asymmetric encoder-decoder architecture with deformable convolutions (DCNs) for HSI-X semantic segmentation. CoMiX is designed to extract, calibrate, and fuse information from HSI and X data. Its pipeline includes an encoder with two parallel and interacting backbones and a lightweight all-multilayer perceptron (ALL-MLP) decoder. The encoder consists of four stages, each incorporating 2D DCN blocks for the X model to accommodate geometric variations and 3D DCN blocks for HSIs to adaptively aggregate spatial-spectral features. Additionally, each stage includes a Cross-Modality Feature enhancement and eXchange (CMFeX) module and a feature fusion module (FFM). CMFeX is designed to exploit spatial-spectral correlations from different modalities to recalibrate and enhance modality-specific and modality-shared features while adaptively exchanging complementary information between them. Outputs from CMFeX are fed into the FFM for fusion and passed to the next stage for further information learning. Finally, the outputs from each FFM are integrated by the ALL-MLP decoder for final prediction. Extensive experiments demonstrate that our CoMiX achieves superior performance and generalizes well to various multimodal recognition tasks. The CoMiX code will be released.

📄 PDF Abstract BibTeX arXiv:2411.09023

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSemantic Segmentation

Similar Papers 제목 키워드 기반

UAVD-Mamba: Deformable Token Fusion Vision Mamba for Multimodal UAV Detection

2025-07-01 · Wei Li, Jiaman Tang, Yang Li, Beihao Xia 외 arxiv

Unmanned Aerial Vehicle (UAV) object detection has been widely used in traffic management, agriculture, emergency rescue, etc. However, it faces significant challenges, including occlusions, small object sizes, and irreg…

Object Detection

CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding

2024-07-04 · Emanuele Vivoli, Marco Bertini, Dimosthenis Karatzas

The comic domain is rapidly advancing with the development of single-page analysis and synthesis models. However, evaluation metrics and datasets lag behind, often limited to small-scale or single-style test sets. We int…

Dialogue Generationobject-detectionObject DetectionSpeaker Identification

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model

2025-05-25 · Alaa Dalaq, Muzammil Behzad

Image segmentation is a fundamental task in computer vision, aimed at partitioning an image into semantically meaningful regions. Referring image segmentation extends this task by using natural language expressions to lo…

cross-modal alignmentImage SegmentationLanguage ModelingLanguage Modelling+3

DAUNet: A Lightweight UNet Variant with Deformable Convolutions and Parameter-Free Attention for Medical Image Segmentation

2025-12-07 · Adnan Munir, Muhammad Shahid Jabbar, Shujaat Khan arxiv

Medical image segmentation plays a pivotal role in automated diagnostic and treatment planning systems. In this work, we present DAUNet, a novel lightweight UNet variant that integrates Deformable V2 Convolutions and Par…

Pulmonary Embolism DetectionMedical Image Segmentation

KLDD: Kalman Filter based Linear Deformable Diffusion Model in Retinal Image Segmentation

2024-09-19 · Zhihao Zhao, Yinzheng Zhao, Junjie Yang, Kai Huang 외

AI-based vascular segmentation is becoming increasingly common in enhancing the screening and treatment of ophthalmic diseases. Deep learning structures based on U-Net have achieved relatively good performance in vascula…

Image SegmentationRetinal Vessel SegmentationSegmentationSemantic Segmentation