paper-with-me

Papers

Touchdown: Natural Language Navigation and Spatial Reasoning in Visual Street Environments

2018-11-29 · CVPR 2019 6 · Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, Yoav Artzi

We study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task. We introduce the Touchdown task and dataset, where an agent must first follow navigation instructions in a real-life visual urban environment, and then identify a location described in natural language to find a hidden object at the goal position. The data contains 9,326 examples of English instructions and spatial descriptions paired with demonstrations. Empirical analysis shows the data presents an open challenge to existing methods, and qualitative linguistic analysis shows that the data displays richer use of spatial reasoning compared to related resources.

📄 PDF Abstract BibTeX arXiv:1811.12354

Code (4)

lil-lab/touchdown 공식 구현 pytorch
VegB/VLN-Transformer pytorch
clic-lab/ciff pytorch
lil-lab/ciff pytorch

Tasks

PositionSpatial ReasoningVision and Language Navigation

Similar Papers 제목 키워드 기반

Retouchdown: Releasing Touchdown on StreetLearn as a Public Resource for Language Grounding Tasks in Street View

2020-11-01 · EMNLP (SpLU) 2020 11 · Harsh Mehta, Yoav Artzi, Jason Baldridge, Eugene Ie 외

The Touchdown dataset (Chen et al., 2019) provides instructions by human annotators for navigation through New York City streets and for resolving spatial descriptions at a given location. To enable the wider research co…

Vision and Language Navigation

Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View

2020-01-10 · Harsh Mehta, Yoav Artzi, Jason Baldridge, Eugene Ie 외

The Touchdown dataset (Chen et al., 2019) provides instructions by human annotators for navigation through New York City streets and for resolving spatial descriptions at a given location. To enable the wider research co…

Vision and Language Navigation

VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation

2024-02-05 · Jialu Li, Aishwarya Padmakumar, Gaurav Sukhatme, Mohit Bansal

Outdoor Vision-and-Language Navigation (VLN) requires an agent to navigate through realistic 3D outdoor environments based on natural language instructions. The performance of existing VLN methods is limited by insuffici…

Language ModelingLanguage ModellingMasked Language ModelingNavigate+1

Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation

2024-08-09 · Huilin Tian, Jingke Meng, Wei-Shi Zheng, Yuan-Ming Li 외

Vision and Language Navigation (VLN) is a challenging task that requires agents to understand instructions and navigate to the destination in a visual environment.One of the key challenges in outdoor VLN is keeping track…

NavigatePositionVision and Language Navigation

SIRI: Spatial Relation Induced Network For Spatial Description Resolution

2020-10-27 · NeurIPS 2020 12 · Peiyao Wang, Weixin Luo, Yanyu Xu, Haojie Li 외

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level rel…

Relation