paper-with-me

홈 › Papers

Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation

2024-09-09 · Muraleekrishna Gopinathan, Martin Masek, Jumana Abu-Khalaf, David Suter

Embodied AI aims to develop robots that can \textit{understand} and execute human language instructions, as well as communicate in natural languages. On this front, we study the task of generating highly detailed navigational instructions for the embodied robots to follow. Although recent studies have demonstrated significant leaps in the generation of step-by-step instructions from sequences of images, the generated instructions lack variety in terms of their referral to objects and landmarks. Existing speaker models learn strategies to evade the evaluation metrics and obtain higher scores even for low-quality sentences. In this work, we propose SAS (Spatially-Aware Speaker), an instruction generator or \textit{Speaker} model that utilises both structural and semantic knowledge of the environment to produce richer instructions. For training, we employ a reward learning method in an adversarial setting to avoid systematic bias introduced by language evaluation metrics. Empirically, our method outperforms existing instruction generation models, evaluated using standard metrics. Our code is available at \url{https://github.com/gmuraleekrishna/SAS}.

📄 PDF Abstract BibTeX arXiv:2409.05583

Code (1)

gmuraleekrishna/sas 공식 구현 pytorch

Tasks

Vision and Language Navigation

Similar Papers 제목 키워드 기반

PASTS: Progress-Aware Spatio-Temporal Transformer Speaker For Vision-and-Language Navigation

2023-05-19 · Liuyi Wang, Chengju Liu, Zongtao He, Shu Li 외

Vision-and-language navigation (VLN) is a crucial but challenging cross-modal navigation task. One powerful technique to enhance the generalization performance in VLN is the use of an independent speaker model to provide…

Data AugmentationVision and Language Navigation

Kefa: A Knowledge Enhanced and Fine-grained Aligned Speaker for Navigation Instruction Generation

2023-07-25 · Haitian Zeng, Xiaohan Wang, Wenguan Wang, Yi Yang

We introduce a novel speaker model \textsc{Kefa} for navigation instruction generation. The existing speaker models in Vision-and-Language Navigation suffer from the large domain gap of vision features between different …

Vision and Language Navigation

FOAM: A Follower-aware Speaker Model For Vision-and-Language Navigation

2022-06-09 · NAACL 2022 7 · Zi-Yi Dou, Nanyun Peng

The speaker-follower models have proven to be effective in vision-and-language navigation, where a speaker model is used to synthesize new instructions to augment the training data for a follower navigation model. Howeve…

Vision and Language Navigation

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation

2026-05-17 · Jingzhi Huang, Junkai Huang, Wenxuan Song, Haoyang Yang 외 arxiv

Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories to directly predict actions, and the zero…

Vision-Language Navigation

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation

2026-04-30 · Pengna Li, Kangyi Wu, Shaoqing Xu, Fang Li 외 arxiv

Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue that adapting VLMs to VLN requires endow…

Vision-Language Navigation