A Joint Speaker-Listener-Reinforcer Model for Referring Expressions
Referring expressions are natural language constructions used to identify particular objects within a scene. In this paper, we propose a unified framework for the tasks of referring expression comprehension and generation. Our model is composed of three modules: speaker, listener, and reinforcer. The speaker generates referring expressions, the listener comprehends referring expressions, and the reinforcer introduces a reward function to guide sampling of more discriminative expressions. The listener-speaker modules are trained jointly in an end-to-end learning framework, allowing the modules to be aware of one another during learning while also benefiting from the discriminative reinforcer's feedback. We demonstrate that this unified framework and training achieves state-of-the-art results for both comprehension and generation on three referring expression datasets. Project and demo page: https://vision.cs.unc.edu/refer
Code (2)
Tasks
Referring ExpressionReferring Expression ComprehensionSimilar Papers 제목 키워드 기반
Learning to Refer to 3D Objects with Natural Language
Human world knowledge is both structured and flexible. When people see an object, they represent it not as a pixel array but as a meaningful arrangement of semantic parts. Moreover, when people refer to an object, they p…
ObjectWorld KnowledgeSpeaking the Language of Your Listener: Audience-Aware Adaptation via Plug-and-Play Theory of Mind
Dialogue participants may have varying levels of knowledge about the topic under discussion. In such cases, it is essential for speakers to adapt their utterances by taking their audience into account. Yet, it is an open…
Language ModelingLanguage ModellingOpen-Ended Question AnsweringText GenerationGrounding Language in Multi-Perspective Referential Communication
We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspecti…
Referring ExpressionReferring expression generationToward Forgetting-Sensitive Referring Expression Generationfor Integrated Robot Architectures
To engage in human-like dialogue, robots require the ability to describe the objects, locations, and people in their environment, a capability known as "Referring Expression Generation." As speakers repeatedly refer to s…
Referring ExpressionReferring expression generationReferring Expressions with Rational Speech Act Framework: A Probabilistic Approach
This paper focuses on a referring expression generation (REG) task in which the aim is to pick out an object in a complex visual scene. One common theoretical approach to this problem is to model the task as a two-agent …
Deep LearningReferring ExpressionReferring expression generation