paper-with-me

홈 › Papers

Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation

2026-05-23 · Biaoyu Ren, Qingsheng Wang, Cun Xu, Dingkang Yang, Wenxuan Wang arxiv

Referring Remote Sensing Image Segmentation (RRSIS) is a situated, task-driven cross-modal task related to the embodied perception paradigm, requiring models to align visual-spatial features with linguistic intentions for precise target perception. Recent research has focused on refining the granularity of textual features and optimizing image-text feature fusion to better guide target feature representations. However, insufficient descriptive granularity and sensitivity to semantic shifts can cause bottlenecks in cross-modal feature fusion. To address these issues, we propose the Image-Conditioned Instance Prompt Network (ICIPNet) with Bilateral Information Fusion, which is designed to alleviate bottlenecks in cross-modal feature fusion. ICIPNet introduces an Image-Conditioned Instance Prompt (ICIP) module to generate self-adaptive visual and semantic representations without external knowledge. The Bilateral Information Fusion (BIF) module enhances feature fusion along the token and channel dimensions. Experiments demonstrate that the proposed ICIPNet outperforms existing RRSIS models.

📄 PDF Abstract BibTeX arXiv:2605.24532

Code (0)

등록된 구현이 없습니다.

Tasks

Image Segmentation

Similar Papers 제목 키워드 기반

Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions

2025-10-26 · Kai Ye, Bowen Liu, Jianghang Lin, Jiayi Ji 외 arxiv

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment instances in remote sensing images according to referring expressions. Unlike Referring Image Segmentation on general images, acquiring high-quality ref…

Referring ExpressionImage Segmentation

Referring Human Pose and Mask Estimation in the Wild

2024-10-27 · Bo Miao, Mingtao Feng, Zijie Wu, Mohammed Bennamoun 외

We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centri…

Decoder

RSRefSeg: Referring Remote Sensing Image Segmentation with Foundation Models

2025-01-12 · Keyan Chen, Jiafan Zhang, Chenyang Liu, Zhengxia Zou 외

Referring remote sensing image segmentation is crucial for achieving fine-grained visual understanding through free-format textual input, enabling enhanced scene and object extraction in remote sensing applications. Curr…

Image SegmentationSegmentationSemantic Segmentation

Customized SAM 2 for Referring Remote Sensing Image Segmentation

2025-03-10 · Fu Rong, Meng Lan, Qian Zhang, Lefei Zhang

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM 2) has shown remarkable performance i…

Image SegmentationSegmentationSemantic Segmentation

Prompt-Driven Referring Image Segmentation with Instance Contrasting

2024-01-01 · CVPR 2024 1 · Chao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang 외

Referring image segmentation (RIS) aims to segment the target referent described by natural language. Recently large-scale pre-trained models e.g. CLIP and SAM have been successfully applied in many downstream tasks …

Contrastive LearningImage SegmentationPrompt LearningSemantic Segmentation