paper-with-me

홈 › Papers

Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction

2018-09-04 · EMNLP 2018 10 · Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, Yoav Artzi

We propose to decompose instruction execution to goal prediction and action generation. We design a model that maps raw visual observations to goals using LINGUNET, a language-conditioned image generation network, and then generates the actions required to complete them. Our model is trained from demonstration only without external resources. To evaluate our approach, we introduce two benchmarks for instruction following: LANI, a navigation task; and CHAI, where an agent executes household instructions. Our evaluation demonstrates the advantages of our model decomposition, and illustrates the challenges posed by our new benchmarks.

📄 PDF Abstract BibTeX arXiv:1809.00786

Code (5)

clic-lab/ciff 공식 구현 pytorch
clic-lab/chalet
clic-lab/drif pytorch
lil-lab/chalet
lil-lab/ciff pytorch

Tasks

Action GenerationConditional Image GenerationImage GenerationInstruction Following

Similar Papers 제목 키워드 기반

ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

2019-12-03 · CVPR 2020 6 · Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk 외

We present ALFRED (Action Learning From Realistic Environments and Directives), a benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks. ALF…

Natural Language Visual Grounding

Scaling Instructable Agents Across Many Simulated Worlds

2024-03-13 · SIMA Team, Maria Abi Raad, Arun Ahuja, Catarina Barros 외

Building embodied AI systems that can follow arbitrary language instructions in any 3D environment is a key challenge for creating general AI. Accomplishing this goal requires learning to ground language in perception an…

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

2025-07-22 · Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang 외 arxiv

Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models …

Robot Manipulation

Following Instructions by Imagining and Reaching Visual Goals

2020-01-25 · John Kanu, Eadom Dessalene, Xiaomin Lin, Cornelia Fermuller 외

While traditional methods for instruction-following typically assume prior linguistic and perceptual knowledge, many recent works in reinforcement learning (RL) have proposed learning policies end-to-end, typically by tr…

Instruction FollowingReinforcement LearningReinforcement Learning (RL)Spatial Reasoning

Are We There Yet? Learning to Localize in Embodied Instruction Following

2021-01-09 · Shane Storks, Qiaozi Gao, Govind Thattai, Gokhan Tur

Embodied instruction following is a challenging problem requiring an agent to infer a sequence of primitive actions to achieve a goal environment state from complex language and visual inputs. Action Learning From Realis…

Instruction Followingobject-detectionObject Detection