Vision-Language Navigation
1개 벤치마크 · 논문 208편 · 이 태스크의 논문 보기 →
Benchmarks
Room2Room
Most implemented
AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN
HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments
The Regretful Agent: Heuristic-Aided Navigation through Progress Estimation
Cross-Lingual Vision-Language Navigation
Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
Papers
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requir…
Vision-Language NavigationLookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory
Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen environments. Recent progress has been largely driven by Multimodal Large Language Models (MLLMs). Existing me…
Vision-Language NavigationIf, Then, Otherwise: Diagnosing Conditional Branching in Vision-Language Navigation
Vision-language navigation agents are often evaluated on their ability to follow route-like instructions toward a fixed goal. Yet, real navigation instructions often depend on observed states of the environment: if a con…
Vision-Language NavigationHumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments
Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: bipedal locomotion imposes physical constraints absent from wheeled agents, humanoid morphologies vary across pla…
Vision-Language NavigationReinforcement LearningAirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN
Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execute feasible 3D motion in complex outdoor…
Vision-Language NavigationTrajectory PredictionSpatial ReasoningWNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly t…
Vision-Language Navigation