Dual Convolutional LSTM Network for Referring Image Segmentation
We consider referring image segmentation. It is a problem at the intersection of computer vision and natural language understanding. Given an input image and a referring expression in the form of a natural language sentence, the goal is to segment the object of interest in the image referred by the linguistic query. To this end, we propose a dual convolutional LSTM (ConvLSTM) network to tackle this problem. Our model consists of an encoder network and a decoder network, where ConvLSTM is used in both encoder and decoder networks to capture spatial and sequential information. The encoder network extracts visual and linguistic features for each word in the expression sentence, and adopts an attention mechanism to focus on words that are more informative in the multimodal interaction. The decoder network integrates the features generated by the encoder network at multiple levels as its input and produces the final precise segmentation mask. Experimental results on four challenging datasets demonstrate that the proposed network achieves superior segmentation performance compared with other state-of-the-art methods.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderImage Segmentationmultimodal interactionNatural Language UnderstandingReferring ExpressionSegmentationSemantic SegmentationSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Recurrent Multimodal Interaction for Referring Image Segmentation
In this paper we are interested in the problem of image segmentation given natural language descriptions, i.e. referring expressions. Existing works tackle this problem by first modeling images and sentences independentl…
Image Segmentationmultimodal interactionSegmentationSemantic SegmentationReferring Image Segmentation by Generative Adversarial Learning
Referring expression is a kind of language expression being used for referring to particular objects. In this paper, we focus on the problem of image segmentation from natural language referring expressions. Existing wor…
Image SegmentationReferring ExpressionReferring Expression SegmentationSegmentation+2Multimodal Referring Segmentation: A Survey
Multimodal referring segmentation aims to segment target objects in visual scenes, such as images, videos, and 3D scenes, based on referring expressions in text or audio format. This task plays a crucial role in practica…
Referring ExpressionSee-Through-Text Grouping for Referring Image Segmentation
Motivated by the conventional grouping techniques to image segmentation, we develop their DNN counterpart to tackle the referring variant. The proposed method is driven by a convolutional-recurrent neural network (ConvRN…
Image Segmentationobject-detectionObject DetectionReferring Expression+4Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network
We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in th…
Referring ExpressionReferring Expression SegmentationSegmentationVideo Segmentation+1