paper-with-me

홈 › Papers

City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs

2025-12-17 · Dwip Dalal, Utkarsh Mishra, Narendra Ahuja, Nebojsa Jojic arxiv

Leveraging multimodal large language models (MLLMs) to develop embodied agents offers significant promise for addressing complex real-world tasks. However, current evaluation benchmarks remain predominantly language-centric or heavily reliant on simulated environments, rarely probing the nuanced, knowledge-intensive reasoning essential for practical, real-world scenarios. To bridge this critical gap, we introduce the task of Sparsely Grounded Visual Navigation, explicitly designed to evaluate the sequential decision-making abilities of MLLMs in challenging, knowledge-intensive real-world environment. We operationalize this task with CityNav, a comprehensive benchmark encompassing four diverse global cities, specifically constructed to assess raw MLLM-driven agents in city navigation. Agents are required to rely solely on visual inputs and internal multimodal reasoning to sequentially navigate 50+ decision points without additional environmental annotations or specialized architectural modifications. Crucially, agents must autonomously achieve localization through interpreting city-specific cues and recognizing landmarks, perform spatial reasoning, and strategically plan and execute routes to their destinations. Through extensive evaluations, we demonstrate that current state-of-the-art MLLMs, reasoning techniques (e.g., GEPA, chain-of-thought, reflection) and competitive baseline PReP significantly underperform in this challenging setting. To address this, we propose Verbalization of Path(VoP), which explicitly grounds the agent's internal reasoning by probing city-scale cognitive maps (key landmarks and directions toward the destination) from the MLLM, substantially enhancing navigation success. Project Webpage: https://dwipddalal.github.io/AgentNav/

📄 PDF Abstract BibTeX arXiv:2512.15933

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningSpatial ReasoningVisual Navigation

Similar Papers 제목 키워드 기반

TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation

2024-03-13 · Dingbang Li, Wenzhou Chen, Xin Lin

Zero-shot navigation is a critical challenge in Vision-Language Navigation (VLN) tasks, where the ability to adapt to unfamiliar instructions and to act in unknown environments is essential. Existing supervised learning-…

Question AnsweringVision-Language Navigation

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

2024-11-26 · CVPR 2025 1 · Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Sujay 외

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods st…

Common Sense ReasoningImitation LearningSpatial ReasoningVisual Navigation

Moving Beyond Navigation with Active Neural SLAM

2022-01-17 · ICLR Track Blog 2022 5 · Anonymous

The ability to effectively harness autonomous control in real-world 3D environments largely depends on learning realistic navigation techniques for embodied agents. Advances to classical robotics call for efficient metho…

Domain Generalizationmotion predictionNavigateScene Understanding

Emergent Braitenberg-style Behaviours for Navigating the ViZDoom `My Way Home' Labyrinth

2024-04-09 · Caleidgh Bayer, Robert J. Smith, Malcolm I. Heywood

The navigation of complex labyrinths with tens of rooms under visual partially observable state is typically addressed using recurrent deep reinforcement learning architectures. In this work, we show that navigation can …

AttributeDeep Reinforcement Learning

Perceive, Reflect, and Plan: Designing LLM Agent for Goal-Directed City Navigation without Instructions

2024-08-08 · Qingbin Zeng, Qinglong Yang, Shunan Dong, Heming Du 외

This paper considers a scenario in city navigation: an AI agent is provided with language descriptions of the goal location with respect to some well-known landmarks; By only observing the scene around, including recogni…

AI AgentNavigate