paper-with-me

Papers

Customized SAM 2 for Referring Remote Sensing Image Segmentation

2025-03-10 · Fu Rong, Meng Lan, Qian Zhang, Lefei Zhang

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM 2) has shown remarkable performance in various segmentation tasks, its application to RRSIS presents several challenges, including understanding the text-described RS scenes and generating effective prompts from text descriptions. To address these issues, we propose RS2-SAM 2, a novel framework that adapts SAM 2 to RRSIS by aligning the adapted RS features and textual features, providing pseudo-mask-based dense prompts, and enforcing boundary constraints. Specifically, we first employ a union encoder to jointly encode the visual and textual inputs, generating aligned visual and text embeddings as well as multimodal class tokens. Then, we design a bidirectional hierarchical fusion module to adapt SAM 2 to RS scenes and align adapted visual features with the visually enhanced text embeddings, improving the model's interpretation of text-described RS scenes. Additionally, a mask prompt generator is introduced to take the visual embeddings and class tokens as input and produce a pseudo-mask as the dense prompt of SAM 2. To further refine segmentation, we introduce a text-guided boundary loss to optimize segmentation boundaries by computing text-weighted gradient differences. Experimental results on several RRSIS benchmarks demonstrate that RS2-SAM 2 achieves state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2503.07266

Code (0)

등록된 구현이 없습니다.

Tasks

Image SegmentationSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
SAM 설명 없음

Similar Papers 제목 키워드 기반

RRSIS: Referring Remote Sensing Image Segmentation

2023-06-14 · Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu

Localizing desired objects from remote sensing images is of great use in practical applications. Referring image segmentation, which aims at segmenting out the objects to which a given expression refers, has been extensi…

BenchmarkingImage SegmentationSegmentationSemantic Segmentation

Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions

2025-10-26 · Kai Ye, Bowen Liu, Jianghang Lin, Jiayi Ji 외 arxiv

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment instances in remote sensing images according to referring expressions. Unlike Referring Image Segmentation on general images, acquiring high-quality ref…

Referring ExpressionImage Segmentation

Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation

2024-09-20 · Sen Lei, Xinyu Xiao, Tianlin Zhang, Heng-Chao Li 외

Given a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixel-wise labels within the imagery. The one of key challenges for this task is to capture disc…

Image SegmentationReferring ExpressionSemantic Segmentation

RSRefSeg: Referring Remote Sensing Image Segmentation with Foundation Models

2025-01-12 · Keyan Chen, Jiafan Zhang, Chenyang Liu, Zhengxia Zou 외

Referring remote sensing image segmentation is crucial for achieving fine-grained visual understanding through free-format textual input, enabling enhanced scene and object extraction in remote sensing applications. Curr…

Image SegmentationSegmentationSemantic Segmentation

AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images

2026-08-09 · Rui Li, Chenxi Duan, Haoyang Yang arxiv

Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote-sensing image segmentation benchmarks accept only written expressio…

Image Segmentation