paper-with-me

홈 › Papers

Grounding Language Attributes to Objects using Bayesian Eigenobjects

2019-05-30 · Vanya Cohen, Benjamin Burchfiel, Thao Nguyen, Nakul Gopalan, Stefanie Tellex, George Konidaris

We develop a system to disambiguate object instances within the same class based on simple physical descriptions. The system takes as input a natural language phrase and a depth image containing a segmented object and predicts how similar the observed object is to the object described by the phrase. Our system is designed to learn from only a small amount of human-labeled language data and generalize to viewpoints not represented in the language-annotated depth image training set. By decoupling 3D shape representation from language representation, this method is able to ground language to novel objects using a small amount of language-annotated depth-data and a larger corpus of unlabeled 3D object meshes, even when these objects are partially observed from unusual viewpoints. Our system is able to disambiguate between novel objects, observed via depth images, based on natural language descriptions. Our method also enables view-point transfer; trained on human-annotated data on a small set of depth images captured from frontal viewpoints, our system successfully predicted object attributes from rear views despite having no such depth images in its training set. Finally, we demonstrate our approach on a Baxter robot, enabling it to pick specific objects based on human-provided natural language descriptions.

📄 PDF Abstract BibTeX arXiv:1905.13153

Code (0)

등록된 구현이 없습니다.

Tasks

3D Shape RepresentationObject

Similar Papers 제목 키워드 기반

Hybrid Bayesian Eigenobjects: Combining Linear Subspace and Deep Network Methods for 3D Robot Vision

2018-06-20 · Benjamin Burchfiel, George Konidaris

We introduce Hybrid Bayesian Eigenobjects (HBEOs), a novel representation for 3D objects designed to allow a robot to jointly estimate the pose, class, and full 3D geometry of a novel object observed from a single viewpo…

3D geometryObjectsubspace methods

Joint Visual Grounding with Language Scene Graphs

2019-06-09 · Daqing Liu, Hanwang Zhang, Zheng-Jun Zha, Meng Wang 외

Visual grounding is a task to ground referring expressions in images, e.g., localize "the white truck in front of the yellow one". To resolve this task fundamentally, the model should first find out the contextual object…

Referring ExpressionVisual Grounding

Grounding Referring Expressions in Images by Variational Context

2017-12-05 · CVPR 2018 6 · Hanwang Zhang, Yulei Niu, Shih-Fu Chang

We focus on grounding (i.e., localizing or linking) referring expressions in images, e.g., "largest elephant standing behind baby elephant". This is a general yet challenging vision-language task since it does not only r…

Multiple Instance LearningReferring Expression

EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding

2022-09-29 · CVPR 2023 1 · Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng 외

3D visual grounding aims to find the object within point clouds mentioned by free-form natural language descriptions with rich semantic cues. However, existing methods either extract the sentence-level features coupling …

3D visual groundingObjectSentenceVisual Grounding

Multi-Attribute Interactions Matter for 3D Visual Grounding

2024-01-01 · CVPR 2024 1 · Can Xu, Yuehui Han, Rui Xu, Le Hui 외

3D visual grounding aims to localize 3D objects described by free-form language sentences. Following the detection-then-matching paradigm existing methods mainly focus on embedding object attributes in unimodal featu…

3D visual groundingAttributeVisual Grounding