paper-with-me

홈 › Papers

Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT Images

2025-03-21 · CVPR 2025 1 · Jie Mei, Chenyu Lin, Yu Qiu, Yaonan Wang, HUI ZHANG, Ziyang Wang, Dong Dai

Lung cancer is a leading cause of cancer-related deaths globally. PET-CT is crucial for imaging lung tumors, providing essential metabolic and anatomical information, while it faces challenges such as poor image quality, motion artifacts, and complex tumor morphology. Deep learning-based models are expected to address these problems, however, existing small-scale and private datasets limit significant performance improvements for these methods. Hence, we introduce a large-scale PET-CT lung tumor segmentation dataset, termed PCLT20K, which comprises 21,930 pairs of PET-CT images from 605 patients. Furthermore, we propose a cross-modal interactive perception network with Mamba (CIPA) for lung tumor segmentation in PET-CT images. Specifically, we design a channel-wise rectification module (CRM) that implements a channel state space block across multi-modal features to learn correlated representations and helps filter out modality-specific noise. A dynamic cross-modality interaction module (DCIM) is designed to effectively integrate position and context information, which employs PET images to learn regional position information and serves as a bridge to assist in modeling the relationships between local features of CT images. Extensive experiments on a comprehensive benchmark demonstrate the effectiveness of our CIPA compared to the current state-of-the-art segmentation methods. We hope our research can provide more exploration opportunities for medical image segmentation. The dataset and code are available at https://github.com/mj129/CIPA.

📄 PDF Abstract BibTeX arXiv:2503.17261

Code (1)

mj129/cipa 공식 구현 pytorch

Tasks

Image SegmentationMambaMedical Image SegmentationPositionSegmentationSemantic SegmentationTumor Segmentation

Methods 이 논문이 사용한 방법론

Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…

Similar Papers 제목 키워드 기반

Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation

2025-10-31 · Elena Mulero Ayllón, Linlin Shen, Pierangelo Veltri, Fabrizia Gelardi 외 arxiv

Accurate lung tumor segmentation is vital for improving diagnosis and treatment planning, and effectively combining anatomical and functional information from PET and CT remains a major challenge. In this study, we propo…

Tumor Segmentation

Multimodal Slice Interaction Network Enhanced by Transfer Learning for Precise Segmentation of Internal Gross Tumor Volume in Lung Cancer PET/CT Imaging

2025-09-26 · Yi Luo, Yike Guo, Hamed Hooshangnejad, Rui Zhang 외 arxiv

Lung cancer remains the leading cause of cancerrelated deaths globally. Accurate delineation of internal gross tumor volume (IGTV) in PET/CT imaging is pivotal for optimal radiation therapy in mobile tumors such as lung …

Transfer Learning

LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks

2025-07-30 · Hui Liu, Chen Jia, Fan Shi, Xu Cheng 외 arxiv

Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interac…

Crack Segmentation

Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery

2025-09-20 · Pengfei Hao, Hongqiu Wang, Shuaibo Li, Zhaohu Xing 외 arxiv

In recent years, Visual Question Localized-Answering in robotic surgery (Surgical-VQLA) has gained significant attention for its potential to assist medical students and junior doctors in understanding surgical scenes. R…

Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion

2026-02-04 · Yixin Zhu, Long Lv, Pingping Zhang, Xuehu Liu 외 arxiv

Multi-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently, some MMIF methods incorporate frequenc…