paper-with-me

Papers

Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training

2020-02-25 · CVPR 2020 6 · Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, Jianfeng Gao

Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often limited. In this paper, we present the first pre-training and fine-tuning paradigm for vision-and-language navigation (VLN) tasks. By training on a large amount of image-text-action triplets in a self-supervised learning manner, the pre-trained model provides generic representations of visual environments and language instructions. It can be easily used as a drop-in for existing VLN frameworks, leading to the proposed agent called Prevalent. It learns more effectively in new tasks and generalizes better in a previously unseen environment. The performance is validated on three VLN tasks. On the Room-to-Room benchmark, our model improves the state-of-the-art from 47% to 51% on success rate weighted by path length. Further, the learned representation is transferable to other VLN tasks. On two recent tasks, vision-and-dialog navigation and "Help, Anna!" the proposed Prevalent leads to significant improvement over existing methods, achieving a new state of the art.

📄 PDF Abstract BibTeX arXiv:2002.10638

Code (1)

weituo12321/PREVALENT 공식 구현

Tasks

NavigateSelf-Supervised LearningVision and Language NavigationVisual Navigation

Similar Papers 제목 키워드 기반

Airbert: In-domain Pretraining for Vision-and-Language Navigation

2021-08-20 · ICCV 2021 10 · Pierre-Louis Guhur, Makarand Tapaswi, ShiZhe Chen, Ivan Laptev 외

Vision-and-language navigation (VLN) aims to enable embodied agents to navigate in realistic environments using natural language instructions. Given the scarcity of domain-specific training data and the high diversity of…

NavigateReferring ExpressionVision and Language Navigation

SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts

2024-12-07 · Gengze Zhou, Yicong Hong, Zun Wang, Chongyang Zhao 외

The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language…

General KnowledgeMixture-of-ExpertsVisual Navigation

Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale

2025-09-29 · Songze Li, Zun Wang, Gengze Zhou, Jialu Li 외 arxiv

Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instructions. Existing methods tend to exclusively …

Vision-Language Navigation

Diagnosing Vision-and-Language Navigation: What Really Matters

2021-03-30 · NAACL 2022 7 · Wanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone 외

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model arc…

DiagnosticObjectVision and Language Navigation

Diagnosing Vision-and-Language Navigation: What Really Matters

2021-12-17 · ACL ARR December 2022 12 · Anonymous

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model arc…

DiagnosticObjectVision and Language Navigation