Referring Expression Generation and Comprehension via Attributes
Referring expression is a kind of language expression that used for referring to particular objects.To make the expression without ambiguation, people often use attributes to describe the particular object. In this paper, we explore the role of attributes by incorporating them into both referring expression generation and comprehension. We first train an attribute learning model from visual objects and their paired descriptions. Then in the generation task, we take the learned attributes as the input into the generation model, thus the expressions are generated driven by both attributes and the previous words. For comprehension, we embed the learned attributes with visual features and semantics into the common space model, then the target object is retrieved based on its ranking distance in the common space. Experimental results on the three standard datasets, RefCOCO, RefCOCO+, and RefCOCOg show significant improvements over the baseline model, demonstrating that our methods are effective for both tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeReferring ExpressionReferring expression generationSimilar Papers 제목 키워드 기반
Comprehension-guided referring expressions
We consider generation and comprehension of natural language referring expression for objects in an image. Unlike generic "image captioning" which lacks natural standard evaluation criteria, quality of a referring expres…
Referring ExpressionReferring expression generationA Joint Speaker-Listener-Reinforcer Model for Referring Expressions
Referring expressions are natural language constructions used to identify particular objects within a scene. In this paper, we propose a unified framework for the tasks of referring expression comprehension and generatio…
Referring ExpressionReferring Expression ComprehensionCo-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages…
Referring ExpressionReferring Expression ComprehensionVideo GroundingLatent Expression Generation for Referring Image Segmentation and Grounding
Visual grounding tasks, such as referring image segmentation (RIS) and referring expression comprehension (REC), aim to localize a target object based on a given textual description. The target object in an image can be …
Generalized Referring Expression SegmentationContrastive LearningImage SegmentationVisual GroundingScene-Text Oriented Reffering Expression Comprehension
Abstract—Referring expression comprehension (REC) aims to identify and locate a specific object in visual scenes referred to by a natural language expression. Existing studies of REC only focus on basic visual attribu…
Object LocalizationReferring ExpressionReferring Expression Comprehension