paper-with-me

Papers

Last-Mile Embodied Visual Navigation

2022-11-21 · Justin Wasserman, Karmesh Yadav, Girish Chowdhary, Abhinav Gupta, Unnat Jain

Realistic long-horizon tasks like image-goal navigation involve exploratory and exploitative phases. Assigned with an image of the goal, an embodied agent must explore to discover the goal, i.e., search efficiently using learned priors. Once the goal is discovered, the agent must accurately calibrate the last-mile of navigation to the goal. As with any robust system, switches between exploratory goal discovery and exploitative last-mile navigation enable better recovery from errors. Following these intuitive guide rails, we propose SLING to improve the performance of existing image-goal navigation systems. Entirely complementing prior methods, we focus on last-mile navigation and leverage the underlying geometric structure of the problem with neural descriptors. With simple but effective switches, we can easily connect SLING with heuristic, reinforcement learning, and neural modular policies. On a standardized image-goal navigation benchmark (Hahn et al. 2021), we improve performance across policies, scenes, and episode complexity, raising the state-of-the-art from 45% to 55% success rate. Beyond photorealistic simulation, we conduct real-robot experiments in three physical scenes and find these improvements to transfer well to real environments.

📄 PDF Abstract BibTeX arXiv:2211.11746

Code (1)

jbwasse2/sling 공식 구현 pytorch

Tasks

Visual Navigation

Similar Papers 제목 키워드 기반

Bridging the Indoor-Outdoor Gap: Vision-Centric Instruction-Guided Embodied Navigation for the Last Meters

2026-02-06 · Yuxiang Zhao, Yirong Yang, Yanqing Zhu, Yanfen Shen 외 arxiv

Embodied navigation holds significant promise for real-world applications such as last-mile delivery. However, most existing approaches are confined to either indoor or outdoor environments and rely heavily on strong ass…

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

2024-11-26 · CVPR 2025 1 · Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Sujay 외

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods st…

Common Sense ReasoningImitation LearningSpatial ReasoningVisual Navigation

UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories

2025-12-10 · Yanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang 외 arxiv

Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dyna…

Spatial ReasoningVisual Navigation

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

2025-11-13 · Xun Huang, Shijia Zhao, Yunxiang Wang, Xin Lu 외 arxiv

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specifi…

CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?

2025-12-18 · Siqi Wang, Chao Liang, Yunfan Gao, Erxin Yu 외 arxiv

Vision-Language Models (VLMs) have made significant progress in explicit instruction-based navigation; however, their ability to interpret implicit human needs (e.g., "I am thirsty") in dynamic urban environments remains…

Spatial Reasoning