Comprehensive Multi-Modal Interactions for Referring Image Segmentation
We investigate Referring Image Segmentation (RIS), which outputs a segmentation map corresponding to the natural language description. Addressing RIS efficiently requires considering the interactions happening across visual and linguistic modalities and the interactions within each modality. Existing methods are limited because they either compute different forms of interactions sequentially (leading to error propagation) or ignore intra-modal interactions. We address this limitation by performing all three interactions simultaneously through a Synchronous Multi-Modal Fusion Module (SFM). Moreover, to produce refined segmentation masks, we propose a novel Hierarchical Cross-Modal Aggregation Module (HCAM), where linguistic features facilitate the exchange of contextual information across the visual hierarchy. We present thorough ablation studies and validate our approach's performance on four benchmark datasets, showing considerable performance gains over the existing state-of-the-art (SOTA) methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Image SegmentationSegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
Comprehensive Multi-Modal Interactions for Referring Image Segmentation
We investigate Referring Image Segmentation (RIS), which outputs a segmentation map corresponding to the natural language description. Addressing RIS efficiently requires considering the interactions happening across vis…
Image SegmentationReferring Expression SegmentationSegmentationSemantic SegmentationImproving Referring Expression Grounding with Cross-modal Attention-guided Erasing
Referring expression grounding aims at locating certain objects or persons in an image with a referring expression, where the key challenge is to comprehend and align various types of information from visual and textual …
Referring ExpressionRefer to Anything with Vision-Language Prompts
Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language an…
BenchmarkingGeneralized Referring Expression SegmentationImage SegmentationReferring Expression+3Multi-Modal Mutual Attention and Iterative Interaction for Referring Image Segmentation
We address the problem of referring image segmentation that aims to generate a mask for the object specified by a natural language expression. Many recent works utilize Transformer to extract features for the target obje…
DecoderImage SegmentationSemantic SegmentationHierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC ext…
Generalized Referring Expression ComprehensionGeneralized Referring Expression SegmentationObject CountingPhrase Grounding+3