paper-with-me

홈 › Papers

XeMap: Contextual Referring in Large-Scale Remote Sensing Environments

2025-04-30 · Yuxi Li, Lu Si, Yujie Hou, Chengaung Liu, Bin Li, Hongjian Fang, Jun Zhang

Advancements in remote sensing (RS) imagery have provided high-resolution detail and vast coverage, yet existing methods, such as image-level captioning/retrieval and object-level detection/segmentation, often fail to capture mid-scale semantic entities essential for interpreting large-scale scenes. To address this, we propose the conteXtual referring Map (XeMap) task, which focuses on contextual, fine-grained localization of text-referred regions in large-scale RS scenes. Unlike traditional approaches, XeMap enables precise mapping of mid-scale semantic entities that are often overlooked in image-level or object-level methods. To achieve this, we introduce XeMap-Network, a novel architecture designed to handle the complexities of pixel-level cross-modal contextual referring mapping in RS. The network includes a fusion layer that applies self- and cross-attention mechanisms to enhance the interaction between text and image embeddings. Furthermore, we propose a Hierarchical Multi-Scale Semantic Alignment (HMSA) module that aligns multiscale visual features with the text semantic vector, enabling precise multimodal matching across large-scale RS imagery. To support XeMap task, we provide a novel, annotated dataset, XeMap-set, specifically tailored for this task, overcoming the lack of XeMap datasets in RS imagery. XeMap-Network is evaluated in a zero-shot setting against state-of-the-art methods, demonstrating superior performance. This highlights its effectiveness in accurately mapping referring regions and providing valuable insights for interpreting large-scale RS environments.

📄 PDF Abstract BibTeX arXiv:2505.00738

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation

2024-09-20 · Sen Lei, Xinyu Xiao, Tianlin Zhang, Heng-Chao Li 외

Given a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixel-wise labels within the imagery. The one of key challenges for this task is to capture disc…

Image SegmentationReferring ExpressionSemantic Segmentation

RRSIS: Referring Remote Sensing Image Segmentation

2023-06-14 · Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu

Localizing desired objects from remote sensing images is of great use in practical applications. Referring image segmentation, which aims at segmenting out the objects to which a given expression refers, has been extensi…

BenchmarkingImage SegmentationSegmentationSemantic Segmentation

Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction Network

2025-08-02 · Jiaxing Yang, Lihe Zhang, Huchuan Lu arxiv

Recently, Referring Remote Sensing Image Segmentation (RRSIS) has aroused wide attention. To handle drastic scale variation of remote targets, existing methods only use the full image as input and nest the saliency-prefe…

Image Segmentation

Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions

2025-10-26 · Kai Ye, Bowen Liu, Jianghang Lin, Jiayi Ji 외 arxiv

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment instances in remote sensing images according to referring expressions. Unlike Referring Image Segmentation on general images, acquiring high-quality ref…

Referring ExpressionImage Segmentation

Multimodal-Aware Fusion Network for Referring Remote Sensing Image Segmentation

2025-03-14 · Leideng Shi, Juan Zhang

Referring remote sensing image segmentation (RRSIS) is a novel visual task in remote sensing images segmentation, which aims to segment objects based on a given text description, with great significance in practical appl…

Image SegmentationSegmentationSemantic Segmentation