paper-with-me

홈 › Papers

Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation

2025-11-02 · Xiangyu Shi, Zerui Li, Yanyuan Qiao, Qi Wu arxiv

Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. However, existing methods often rely on panoramic observations and two-stage pipelines involving waypoint predictors, which introduce significant latency and limit real-world applicability. In this work, we propose Fast-SmartWay, an end-to-end zero-shot VLN-CE framework that eliminates the need for panoramic views and waypoint predictors. Our approach uses only three frontal RGB-D images combined with natural language instructions, enabling MLLMs to directly predict actions. To enhance decision robustness, we introduce an Uncertainty-Aware Reasoning module that integrates (i) a Disambiguation Module for avoiding local optima, and (ii) a Future-Past Bidirectional Reasoning mechanism for globally coherent planning. Experiments on both simulated and real-robot environments demonstrate that our method significantly reduces per-step latency while achieving competitive or superior performance compared to panoramic-view baselines. These results demonstrate the practicality and effectiveness of Fast-SmartWay for real-world zero-shot embodied navigation.

📄 PDF Abstract BibTeX arXiv:2511.00933

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE

2026-06-30 · Or Hirschorn, Aaron Olender, Eli Alshan, Ianir Ideses 외 hf

We present a zero-shot, training-free and optimization-free framework for generating 360 panoramic images and videos by directly injecting spherical priors into pre-trained diffusion transformers. Existing methods either…

VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation

2026-03-19 · Jiayi Yuan, Haobo Jiang, De Wen Soh, Na Zhao arxiv

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panor…

3D ReconstructionDepth Estimation

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

2025-03-13 · Xiangyu Shi, Zerui Li, Wenqi Lyu, Jiatong Xia 외

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach…

Language ModelingLanguage ModellingLarge Language ModelVision and Language Navigation

DA$^{2}$: Depth Anything in Any Direction

2025-09-30 · Haodong Li, Wangguangdong Zheng, Jing He, Yuhao Liu 외 arxiv

Panorama has a full FoV (360$^\circ\times$180$^\circ$), offering a more complete visual description than perspective images. Thanks to this characteristic, panoramic depth estimation is gaining increasing traction in 3D …

Zero-shot GeneralizationDepth Estimation

P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation

2026-05-19 · Kai Sheng, Liuyi Wang, Haojie Dai, Jinlong Li 외 arxiv

Vision-and-language navigation (VLN) requires an embodied agent to ground natural-language instructions into executable navigation actions in unseen environments. Existing zero-shot methods typically rely on additional w…