paper-with-me

Papers

Learning to Refer to 3D Objects with Natural Language

2019-05-01 · ICLR 2019 5 · Panos Achlioptas, Judy E. Fan, Robert X. D. Hawkins, Noah D. Goodman, Leo Guibas

Human world knowledge is both structured and flexible. When people see an object, they represent it not as a pixel array but as a meaningful arrangement of semantic parts. Moreover, when people refer to an object, they provide descriptions that are not merely true but also relevant in the current context. Here, we combine these two observations in order to learn fine-grained correspondences between language and contextually relevant geometric properties of 3D objects. To do this, we employed an interactive communication task with human participants to construct a large dataset containing natural utterances referring to 3D objects from ShapeNet in a wide variety of contexts. Using this dataset, we developed neural listener and speaker models with strong capacity for generalization. By performing targeted lesions of visual and linguistic input, we discovered that the neural listener depends heavily on part-related words and associates these words correctly with the corresponding geometric properties of objects, suggesting that it has learned task-relevant structure linking the two input modalities. We further show that a neural speaker that is listener-aware' --- that plans its utterances according to how an imagined listener would interpret its words in context --- produces more discriminative referring expressions than an listener-unaware' speaker, as measured by human performance in identifying the correct object.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectWorld Knowledge

Similar Papers 제목 키워드 기반

Grounding Spatio-Semantic Referring Expressions for Human-Robot Interaction

2017-07-18 · Mohit Shridhar, David Hsu

The human language is one of the most natural interfaces for humans to interact with robots. This paper presents a robot system that retrieves everyday objects with unconstrained natural language descriptions. A core iss…

Object

Temporal Collection and Distribution for Referring Video Object Segmentation

2023-09-07 · ICCV 2023 1 · Jiajin Tang, Ge Zheng, Sibei Yang

Referring video object segmentation aims to segment a referent throughout a video sequence according to a natural language expression. It requires aligning the natural language expression with the objects' motions and th…

ObjectReferring Video Object SegmentationSemantic SegmentationVideo Object Segmentation+1

Modeling Context in Referring Expressions

2016-07-31 · Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg 외

Humans refer to objects in their environments all the time, especially in dialogue with other people. We explore generating and comprehending natural language referring expressions for objects in images. In particular, w…

Referring ExpressionReferring expression generationText Generation

Multi3DRefer: Grounding Text Description to Multiple 3D Objects

2023-09-11 · ICCV 2023 1 · Yiming Zhang, ZeMing Gong, Angel X. Chang

We introduce the task of localizing a flexible number of objects in real-world 3D scenes using natural language descriptions. Existing 3D visual grounding tasks focus on localizing a unique object given a text descriptio…

3D visual groundingContrastive LearningObjectObject Rearrangement+3

Language Grounding with 3D Objects

2021-07-26 · Jesse Thomason, Mohit Shridhar, Yonatan Bisk, Chris Paxton 외

Seemingly simple natural language requests to a robot are generally underspecified, for example "Can you bring me the wireless mouse?" Flat images of candidate mice may not provide the discriminative information needed f…