paper-with-me

Papers

Transferable Representation Learning in Vision-and-Language Navigation

2019-08-09 · ICCV 2019 10 · Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Magalhaes, Jason Baldridge, Eugene Ie

Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic environments to achieve navigation goals. The overall task requires competence in several perception problems: successful agents combine spatio-temporal, vision and language understanding to produce appropriate action sequences. Our approach adapts pre-trained vision and language representations to relevant in-domain tasks making them more effective for VLN. Specifically, the representations are adapted to solve both a cross-modal sequence alignment and sequence coherence task. In the sequence alignment task, the model determines whether an instruction corresponds to a sequence of visual frames. In the sequence coherence task, the model determines whether the perceptual sequences are predictive sequentially in the instruction-conditioned latent space. By transferring the domain-adapted representations, we improve competitive agents in R2R as measured by the success rate weighted by path length (SPL) metric.

📄 PDF Abstract BibTeX arXiv:1908.03409

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningVision and Language Navigation

Similar Papers 제목 키워드 기반

Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training

2020-02-25 · CVPR 2020 6 · Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 외

Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often lim…

NavigateSelf-Supervised LearningVision and Language NavigationVisual Navigation

SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation

2026-07-16 · Yi Wu, Junjie An, Xiao Liu, Yiqun Zhou 외 arxiv

In goal-directed embodied navigation, where an agent must locate a specified target in an unseen environment, 3D scene understanding and navigation reasoning must work in concert. Current approaches transmit 3D scene inf…

Scene Understanding

Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale

2025-09-29 · Songze Li, Zun Wang, Gengze Zhou, Jialu Li 외 arxiv

Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instructions. Existing methods tend to exclusively …

Vision-Language Navigation

Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation

2019-11-18 · CVPR 2020 6 · Juncheng Li, Xin Wang, Siliang Tang, Haizhou Shi 외

Visual navigation is a task of training an embodied agent by intelligently navigating to a target object (e.g., television) using only visual observations. A key challenge for current deep reinforcement learning models l…

Deep Reinforcement LearningObjectreinforcement-learningReinforcement Learning+3

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

2026-05-28 · Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye 외 arxiv

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and r…

Trajectory PredictionSpatial ReasoningVisual Grounding