Explore the Potential Performance of Vision-and-Language Navigation Model: a Snapshot Ensemble Method
Vision-and-Language Navigation (VLN) is a challenging task in the field of artificial intelligence. Although massive progress has been made in this task over the past few years attributed to breakthroughs in deep vision and language models, it remains tough to build VLN models that can generalize as well as humans. In this paper, we provide a new perspective to improve VLN models. Based on our discovery that snapshots of the same VLN model behave significantly differently even when their success rates are relatively the same, we propose a snapshot-based ensemble solution that leverages predictions among multiple snapshots. Constructed on the snapshots of the existing state-of-the-art (SOTA) model $\circlearrowright$BERT and our past-action-aware modification, our proposed ensemble achieves the new SOTA performance in the R2R dataset challenge in Navigation Error (NE) and Success weighted by Path Length (SPL).
Code (0)
등록된 구현이 없습니다.
Tasks
Vision and Language NavigationSimilar Papers 제목 키워드 기반
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
This paper explores the potential of Large Language Models(LLMs) in zero-shot anomaly detection for safe visual navigation. With the assistance of the state-of-the-art real-time open-world object detection model Yolo-Wor…
Anomaly Detectionobject-detectionObject DetectionOpen-vocabulary object detection+5Explore the Potential Performance of Vision-and-Language Navigation Model: a Snapshot Ensemble Method
Given an instruction in a natural language, the vision-and-language navigation (VLN) task requires a navigation model to match the instruction to its visual surroundings and then move to the correct destination. It has b…
Vision and Language NavigationLangNav: Language as a Perceptual Representation for Navigation
We explore the use of language as a perceptual representation for vision-and-language navigation (VLN), with a focus on low-data settings. Our approach uses off-the-shelf vision systems for image captioning and object de…
Image CaptioningLanguage ModelingLanguage ModellingLarge Language Model+3Sim-2-Sim Transfer for Vision-and-Language Navigation in Continuous Environments
Recent work in Vision-and-Language Navigation (VLN) has presented two environmental paradigms with differing realism -- the standard VLN setting built on topological environments where navigation is abstracted away, and …
NavigateVision and Language NavigationImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
Visual navigation is a fundamental capability for autonomous home-assistance robots, enabling long-horizon tasks such as object search. While recent methods have leveraged Large Language Models (LLMs) to incorporate comm…
Spatial ReasoningVisual Navigation