paper-with-me

홈 › Papers

Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution

2026-05-12 · Chen Wu, Ling Wang, Zhuoran Zheng, Xiangyu Chen, Jingyuan Xia, Weidong Jiang, Jiantao Zhou arxiv

Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality independently or rely on computationally expensive attention mechanisms with quadratic complexity, hindering the establishment of efficient and semantically interactive joint representations. In this paper, we observe that feature maps from different modalities exhibit semantic-level correlations during feature extraction. This motivates us to develop a more flexible approach enabling dense, semantically-aware deep interactions between modalities. To this end, we propose a novel GDSR framework centered around the Interactive State Space Model. Specifically, we design a cross-modal local scanning mechanism that enables fine-grained semantic interactions between RGB and depth features. Leveraging the Mamba architecture, our framework achieves global modeling with linear complexity. Furthermore, a cross-modal matching transform module is introduced to enhance interactive modeling quality by utilizing representative features from both modalities. Extensive experiments demonstrate competitive performance against state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2605.11934

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

2024-07-18 · Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 외

Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, …

cross-modal alignmentCross-Modal RetrievalSound Source Localization

EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents

2026-05-02 · Sai Ma, Zhuang Li, Sichao Li, Xinyue Xu 외 arxiv

Earth Observation (EO) analysis is inherently interactive: resolving uncertainty often requires expanding the region of interest, retrieving historical observations, and switching across sensors such as optical and Synth…

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

Cross-modal Retrieval with Improved Graph Convolution

2023-03-07 · Computer Engineering and Applications 2023 3 · ZHANG Hongtu, HUA Chunjian, JIANG Yi, YU Jianfeng 외

Aiming at the problem that existing image text cross-modal retrieval is difficult to fully exploit the local consistency in the mode in the common subspace, a cross-modal retrieval method based on improved graph convolu…

Cross-Modal RetrievalRepresentation LearningRetrievalSentence

Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT Images

2025-03-21 · CVPR 2025 1 · Jie Mei, Chenyu Lin, Yu Qiu, Yaonan Wang 외

Lung cancer is a leading cause of cancer-related deaths globally. PET-CT is crucial for imaging lung tumors, providing essential metabolic and anatomical information, while it faces challenges such as poor image quality,…

Image SegmentationMambaMedical Image SegmentationPosition+3