paper-with-me

홈 › Papers

Beyond One-to-One: Rethinking the Referring Image Segmentation

2023-08-26 · ICCV 2023 1 · Yutao Hu, Qixiong Wang, Wenqi Shao, Enze Xie, Zhenguo Li, Jungong Han, Ping Luo

Referring image segmentation aims to segment the target object referred by a natural language expression. However, previous methods rely on the strong assumption that one sentence must describe one target in the image, which is often not the case in real-world applications. As a result, such methods fail when the expressions refer to either no objects or multiple objects. In this paper, we address this issue from two perspectives. First, we propose a Dual Multi-Modal Interaction (DMMI) Network, which contains two decoder branches and enables information flow in two directions. In the text-to-image decoder, text embedding is utilized to query the visual feature and localize the corresponding target. Meanwhile, the image-to-text decoder is implemented to reconstruct the erased entity-phrase conditioned on the visual feature. In this way, visual features are encouraged to contain the critical semantic information about target entity, which supports the accurate segmentation in the text-to-image decoder in turn. Secondly, we collect a new challenging but realistic dataset called Ref-ZOM, which includes image-text pairs under different settings. Extensive experiments demonstrate our method achieves state-of-the-art performance on different datasets, and the Ref-ZOM-trained model performs well on various types of text inputs. Codes and datasets are available at https://github.com/toggle1995/RIS-DMMI.

📄 PDF Abstract BibTeX arXiv:2308.13853

Code (1)

toggle1995/ris-dmmi 공식 구현 pytorch

Tasks

DecoderImage SegmentationImage to textSemantic SegmentationSentence

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Temporal Prompting Matters: Rethinking Referring Video Object Segmentation

2025-10-08 · Ci-Siang Lin, Min-Hung Chen, I-Jieh Liu, Chien-Yi Wang 외 arxiv

Referring Video Object Segmentation (RVOS) aims to segment the object referred to by the query sentence in the video. Most existing methods require end-to-end training with dense mask annotations, which could be computat…

Referring Video Object Segmentation

RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation

2025-09-16 · Siju Ma, Changsiyu Gong, Xiaofeng Fan, Yong Ma 외 arxiv

Text-driven infrared and visible image fusion has gained attention for enabling natural language to guide the fusion process. However, existing methods lack a goal-aligned task to supervise and evaluate how effectively t…

Referring ExpressionImage Segmentation

Rethinking Referring Object Removal

2024-03-14 · Xiangtian Xue, Jiasong Wu, Youyong Kong, Lotfi Senhadji 외

Referring object removal refers to removing the specific object in an image referred by natural language expressions and filling the missing region with reasonable semantics. To address this task, we construct the ComCOC…

Object

Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation

2021-06-02 · Chen Liang, Yu Wu, Tianfei Zhou, Wenguan Wang 외

Referring video object segmentation (RVOS) aims to segment video objects with the guidance of natural language reference. Previous methods typically tackle RVOS through directly grounding linguistic reference over the im…

ObjectOne-shot visual object segmentationReferring Video Object SegmentationSemantic Segmentation+2

Advancing Referring Expression Segmentation Beyond Single Image

2023-05-21 · ICCV 2023 1 · Yixuan Wu, Zhao Zhang, Xie Chi, Feng Zhu 외

Referring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expression. However, in broader real-world s…

Co-Salient Object DetectionObjectobject-detectionObject Detection+3