Vision and Language Navigation
5개 벤치마크 · 논문 224편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
How Much Can CLIP Benefit Vision-and-Language Tasks?
Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View
Papers
AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
Aerial Vision-and-Language Navigation (VLN) is an emerging task that enables Unmanned Aerial Vehicles (UAVs) to navigate outdoor environments using natural language instructions and visual cues. However, due to the exten…
Vision and Language NavigationRethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
Recent Vision-and-Language Navigation (VLN) advancements are promising, but their idealized assumptions about robot movement and control fail to reflect physically embodied deployment challenges. To bridge this gap, we i…
Large Language ModelVision and Language NavigationNavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often strugg…
Decision MakingVision and Language NavigationGrounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human instructions while navigating complex environments. Two key bottlenecks …
Language ModelingLanguage ModellingLarge Language ModelVision and Language Navigation+1A Navigation Framework Utilizing Vision-Language Models
Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, unfamiliar environments. Recent advances i…
NavigatePrompt EngineeringVision and Language NavigationDisrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
We present Adversarial Object Fusion (AdvOF), a novel attack framework targeting vision-and-language navigation (VLN) agents in service-oriented environments by generating adversarial 3D objects. While foundational model…
Language ModelingLanguage ModellingObjectService Composition+1