Few-shot Object Grounding and Mapping for Natural Language Robot Instruction Following
We study the problem of learning a robot policy to follow natural language instructions that can be easily extended to reason about new objects. We introduce a few-shot language-conditioned object grounding method trained from augmented reality data that uses exemplars to identify objects and align them to their mentions in instructions. We present a learned map representation that encodes object locations and their instructed use, and construct it from our few-shot grounding output. We integrate this mapping approach into an instruction-following policy, thereby allowing it to reason about previously unseen objects at test-time by simply adding exemplars. We evaluate on the task of learning to map raw observations and instructions to continuous control of a physical quadcopter. Our approach significantly outperforms the prior state of the art in the presence of new objects, even when the prior approach observes all objects during training.
Code (1)
Tasks
continuous-controlContinuous ControlInstruction FollowingSimilar Papers 제목 키워드 기반
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic repre…
Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems
3D visual grounding (3DVG) aims to locate objects in a 3D scene with natural language descriptions. Supervised methods have achieved decent accuracy, but have a closed vocabulary and limited language understanding abilit…
3D visual groundingNegationVisual GroundingEnd-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding
Natural language spatial video grounding aims to detect the relevant objects in video frames with descriptive sentences as the query. In spite of the great advances, most existing methods rely on dense video frame annota…
DescriptiveRepresentation LearningVideo GroundingZero-Shot Grounding of Objects from Natural Language Queries
A phrase grounding system localizes a particular object in an image referred to by a natural language query. In previous work, the phrases were restricted to have nouns that were encountered in training, we extend the ta…
Natural Language Queriesobject-detectionObject DetectionPhrase GroundingVLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
3D visual grounding is crucial for robots, requiring integration of natural language and 3D scene understanding. Traditional methods depending on supervised learning with 3D point clouds are limited by scarce datasets. R…
3D geometry3D visual groundingObjectScene Understanding+1