paper-with-me

홈 › Papers

A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning

2022-10-06 · CVPR 2023 1 · Aishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, Zarana Parekh

Recent studies in Vision-and-Language Navigation (VLN) train RL agents to execute natural-language navigation instructions in photorealistic environments, as a step towards robots that can follow human instructions. However, given the scarcity of human instruction data and limited diversity in the training environments, these agents still struggle with complex language grounding and spatial language understanding. Pretraining on large text and image-text datasets from the web has been extensively explored but the improvements are limited. We investigate large-scale augmentation with synthetic instructions. We take 500+ indoor environments captured in densely-sampled 360 degree panoramas, construct navigation trajectories through these panoramas, and generate a visually-grounded instruction for each trajectory using Marky, a high-quality multilingual navigation instruction generator. We also synthesize image observations from novel viewpoints using an image-to-image GAN. The resulting dataset of 4.2M instruction-trajectory pairs is two orders of magnitude larger than existing human-annotated datasets, and contains a wider variety of environments and viewpoints. To efficiently leverage data at this scale, we train a simple transformer agent with imitation learning. On the challenging RxR dataset, our approach outperforms all existing RL agents, improving the state-of-the-art NDTW from 71.1 to 79.1 in seen environments, and from 64.6 to 66.8 in unseen test environments. Our work points to a new path to improving instruction-following agents, emphasizing large-scale imitation learning and the development of synthetic instruction generation capabilities.

📄 PDF Abstract BibTeX arXiv:2210.03112

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningInstruction FollowingVision and Language Navigation

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics

2026-02-11 · Mengzhou Wu, Yuzhe Guo, Yuan Cao, Haochuan Lu 외 arxiv

Scaling generalist GUI agents is hindered by the data scalability bottleneck of expensive human demonstrations and the "distillation ceiling" of synthetic teacher supervision. To transcend these limitations, we propose U…

OpenVLN: Open-world Aerial Vision-Language Navigation

2025-11-09 · Peican Lin, Gan Sun, Chenxi Liu, Fazeng Li 외 arxiv

Vision-language models (VLMs) have been widely-applied in ground-based vision-language navigation (VLN). However, the vast complexity of outdoor aerial environments compounds data acquisition challenges and imposes long-…

Vision-Language NavigationReinforcement LearningTrajectory Planning

What Limits Vision-and-Language Navigation ?

2026-05-13 · Yunheng Wang, Yuetong Fang, Taowen Wang, Lusong Li 외 arxiv

Vision-and-Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning from simulation to real-world deployment, p…

NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

2023-05-26 · Gengze Zhou, Yicong Hong, Qi Wu

Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of trai…

Instruction FollowingVision and Language NavigationVisual Navigation

Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments

2021-09-30 · EMNLP 2021 11 · Sonia Raychaudhuri, Saim Wani, Shivansh Patel, Unnat Jain 외

In the Vision-and-Language Navigation (VLN) task an embodied agent navigates a 3D environment, following natural language instructions. A challenge in this task is how to handle 'off the path' scenarios where an agent ve…

Vision and Language Navigation