paper-with-me

홈 › Papers

Dual Convolutional LSTM Network for Referring Image Segmentation

2020-01-30 · Linwei Ye, Zhi Liu, Yang Wang

We consider referring image segmentation. It is a problem at the intersection of computer vision and natural language understanding. Given an input image and a referring expression in the form of a natural language sentence, the goal is to segment the object of interest in the image referred by the linguistic query. To this end, we propose a dual convolutional LSTM (ConvLSTM) network to tackle this problem. Our model consists of an encoder network and a decoder network, where ConvLSTM is used in both encoder and decoder networks to capture spatial and sequential information. The encoder network extracts visual and linguistic features for each word in the expression sentence, and adopts an attention mechanism to focus on words that are more informative in the multimodal interaction. The decoder network integrates the features generated by the encoder network at multiple levels as its input and produces the final precise segmentation mask. Experimental results on four challenging datasets demonstrate that the proposed network achieves superior segmentation performance compared with other state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2001.11561

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage Segmentationmultimodal interactionNatural Language UnderstandingReferring ExpressionSegmentationSemantic SegmentationSentence

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
ConvLSTM ConvLSTM is a type of recurrent neural network for spatio-temporal prediction that has convolutional structures in both the input-to-state and state-to-state transitions. The…
Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Recurrent Multimodal Interaction for Referring Image Segmentation

2017-03-23 · ICCV 2017 10 · Chenxi Liu, Zhe Lin, Xiaohui Shen, Jimei Yang 외

In this paper we are interested in the problem of image segmentation given natural language descriptions, i.e. referring expressions. Existing works tackle this problem by first modeling images and sentences independentl…

Image Segmentationmultimodal interactionSegmentationSemantic Segmentation

Referring Image Segmentation by Generative Adversarial Learning

2020-04-20 · IEEE 2020 4 · Shuang Qiu, Yao Zhao, Jianbo Jiao, Yunchao Wei 외

Referring expression is a kind of language expression being used for referring to particular objects. In this paper, we focus on the problem of image segmentation from natural language referring expressions. Existing wor…

Image SegmentationReferring ExpressionReferring Expression SegmentationSegmentation+2

Multimodal Referring Segmentation: A Survey

2025-08-01 · Henghui Ding, Song Tang, Shuting He, Chang Liu 외 arxiv

Multimodal referring segmentation aims to segment target objects in visual scenes, such as images, videos, and 3D scenes, based on referring expressions in text or audio format. This task plays a crucial role in practica…

Referring Expression

See-Through-Text Grouping for Referring Image Segmentation

2019-10-01 · ICCV 2019 10 · Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen 외

Motivated by the conventional grouping techniques to image segmentation, we develop their DNN counterpart to tackle the referring variant. The proposed method is driven by a convolutional-recurrent neural network (ConvRN…

Image Segmentationobject-detectionObject DetectionReferring Expression+4

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

2021-02-09 · Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang 외

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in th…

Referring ExpressionReferring Expression SegmentationSegmentationVideo Segmentation+1