paper-with-me

홈 › Papers

Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation

2025-10-31 · Elena Mulero Ayllón, Linlin Shen, Pierangelo Veltri, Fabrizia Gelardi, Arturo Chiti, Paolo Soda, Matteo Tortora arxiv

Accurate lung tumor segmentation is vital for improving diagnosis and treatment planning, and effectively combining anatomical and functional information from PET and CT remains a major challenge. In this study, we propose vMambaX, a lightweight multimodal framework integrating PET and CT scan images through a Context-Gated Cross-Modal Perception Module (CGM). Built on the Visual Mamba architecture, vMambaX adaptively enhances inter-modality feature interaction, emphasizing informative regions while suppressing noise. Evaluated on the PCLT20K dataset, the model outperforms baseline models while maintaining lower computational complexity. These results highlight the effectiveness of adaptive cross-modal gating for multimodal tumor segmentation and demonstrate the potential of vMambaX as an efficient and scalable framework for advanced lung cancer analysis. The code is available at https://github.com/arco-group/vMambaX.

📄 PDF Abstract BibTeX arXiv:2510.27508

Code (0)

등록된 구현이 없습니다.

Tasks

Tumor Segmentation

Similar Papers 제목 키워드 기반

Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions

2026-08-20 · Mohammad Arif Ul Alam arxiv

Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur. Although sonar provides complementary i…

Contextual Object Detection with Multimodal Large Language Models

2023-05-29 · Yuhang Zang, Wei Li, Jun Han, Kaiyang Zhou 외

Recent Multimodal Large Language Models (MLLMs) are remarkable in vision-language tasks, such as image captioning and question answering, but lack the essential perception ability, i.e., object detection. In this work, w…

Cloze TestDecoderImage CaptioningImage Segmentation+5

MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception

2024-06-22 · Guanqun Wang, Xinyu Wei, Jiaming Liu, Ray Zhang 외

In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual perception models have made significant stride…

Common Sense ReasoningLanguage ModellingLarge Language ModelMultimodal Large Language Model+4

Multimodal perception for dexterous manipulation

2021-12-28 · Guanqun Cao, Shan Luo

Humans usually perceive the world in a multimodal way that vision, touch, sound are utilised to understand surroundings from various dimensions. These senses are combined together to achieve a synergistic effect where th…

3D ReconstructionFrictionTranslation

M$^3$-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering

2026-03-09 · Peijin Xie, Zhen Xu, Bingquan Liu, Baoxun Wang arxiv

Multimodal large language models have recently shown promising progress in visual mathematical reasoning. However, their performance is often limited by a critical yet underexplored bottleneck: inaccurate visual percepti…

Mathematical ReasoningMultimodal ReasoningPrompt Engineering