paper-with-me

Vision-Language Navigation

1개 벤치마크 · 논문 208편 · 이 태스크의 논문 보기 →

Benchmarks

Room2Room

결과 3개

Most implemented

Cross-Lingual Vision-Language Navigation

2019-10-24 · 구현 2개

Papers

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

2026-09-08 · Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao 외 hf

We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requir…

Vision-Language Navigation

LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory

2026-09-02 · Kun-Yang Yu, Yingzhe Li, Hongyu Xu, Shi-Yu Tian 외 arxiv

Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen environments. Recent progress has been largely driven by Multimodal Large Language Models (MLLMs). Existing me…

Vision-Language Navigation

If, Then, Otherwise: Diagnosing Conditional Branching in Vision-Language Navigation

2026-08-18 · Seoyoung Lee, Neel P. Bhatt, Pranay Samineni, Cong Liu 외 arxiv

Vision-language navigation agents are often evaluated on their ability to follow route-like instructions toward a fixed goal. Yet, real navigation instructions often depend on observed states of the environment: if a con…

Vision-Language Navigation

HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments

2026-08-13 · Quan-Dung Pham, Anh Dao, The-Anh Nguyen, Minh Nguyen-Dinh 외 arxiv

Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: bipedal locomotion imposes physical constraints absent from wheeled agents, humanoid morphologies vary across pla…

Vision-Language NavigationReinforcement Learning

AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

2026-08-13 · Yutong Liu, Xiaojie Li, Mingzhu Xu, Jianlong Wu arxiv

Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execute feasible 3D motion in complex outdoor…

Vision-Language NavigationTrajectory PredictionSpatial Reasoning

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

2026-08-07 · Yuehao Huang, Yunzi Wu, Xiaotao Zhang, Xinhai Li 외 arxiv

Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly t…

Vision-Language Navigation

전체 208편 보기 →