paper-with-me

Papers

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

2024-11-26 · CVPR 2025 1 · Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Sujay, Zhicheng Yang, Juexiao Zhang, John Abanes, Jing Zhang, Chen Feng

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous agents like last-mile delivery robots. To overcome these obstacles, we propose a scalable, data-driven approach for human-like urban navigation by training agents on thousands of hours of in-the-wild city walking and driving videos sourced from the web. We introduce a simple and scalable data processing pipeline that extracts action supervision from these videos, enabling large-scale imitation learning without costly annotations. Our model learns sophisticated navigation policies to handle diverse challenges and critical scenarios. Experimental results show that training on large-scale, diverse datasets significantly enhances navigation performance, surpassing current methods. This work shows the potential of using abundant online video data to develop robust navigation policies for embodied agents in dynamic urban settings. Project homepage is at https://ai4ce.github.io/CityWalker/.

📄 PDF Abstract BibTeX arXiv:2411.17820

Code (1)

ai4ce/CityWalker 공식 구현 pytorch

Tasks

Common Sense ReasoningImitation LearningSpatial ReasoningVisual Navigation

Similar Papers 제목 키워드 기반

Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos

2026-08-17 · Bingyi Xia, Han Bao, Zhewei Chen, Hanjing Ye 외 arxiv

Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific data collection and the limited coverage of rare yet safety-critical scenarios. To address these challenges, we…

UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories

2025-12-10 · Yanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang 외 arxiv

Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dyna…

Spatial ReasoningVisual Navigation

Learning to Generate Diverse Pedestrian Movements from Web Videos with Noisy Labels

2024-10-10 · Zhizheng Liu, Joe Lin, Wayne Wu, Bolei Zhou

Understanding and modeling pedestrian movements in the real world is crucial for applications like motion forecasting and scene simulation. Many factors influence pedestrian movements, such as scene context, individual c…

Motion ForecastingZero-shot Generalization

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

2026-08-09 · Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa 외 hf

We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack su…

Spatial Reasoning

UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos

2025-10-16 · Mingxuan Liu, Honglin He, Elisa Ricci, Wayne Wu 외 arxiv

Urban embodied AI agents, ranging from delivery robots to quadrupeds, are increasingly populating our cities, navigating chaotic streets to provide last-mile connectivity. Training such agents requires diverse, high-fide…