paper-with-me

홈 › Papers

Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation

2026-09-15 · Hojin Lee, Sizhe Lester Li, Maximilian Hilger, Susie Lu, Achim J. Lilienthal, Vincent Sitzmann, Daniel A. Duecker arxiv

Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guidance and recover geometric waypoints through scene reconstruction, leaving longer-horizon planning and precise video-to-action translation less explored. We present CueNav, a video model-based navigation framework combining visual cue guided video planning with an embodiment-specific Inverse-Dynamics Model (IDM). As visual cues, we use a Bird's-Eye View (BEV) map to convey global task context and retain part of the robot body in the egocentric observation to expose embodiment context. These cues guide the video planner, while the IDM translates dense flow fields extracted from the video plan into robot actions. With the visual cue encoding global task context, CueNav achieves nearly 2x higher success in maze navigation than planning without the cue. The body-aware view with the IDM enables precise navigation with 70% success in a narrow passage where comparison methods largely fail to complete the task. We further demonstrate zero-shot semantic-conditioned navigation and deployment of the same video planner across different robot platforms. Our results show that visual cue-guided video planning with embodiment-specific action grounding paves the way toward a generalizable navigation framework for longer-horizon planning and embodiment-aware control. Additional results and code are available on our project website: https://cuenav.github.io.

📄 PDF Abstract BibTeX arXiv:2609.16737

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Navigation

Similar Papers 제목 키워드 기반

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

2026-07-09 · Qi Lyu, Baicheng Liu, Xudong Wang, Jiahua Dong 외 arxiv

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning w…

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

2025-11-24 · Ziqi Ni, Yuanzhi Liang, Rui Li, Yi Zhou 외 arxiv

Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, exi…

Reinforcement Learning

Seeing What You're Told: Sentence-Guided Activity Recognition In Video

2013-08-19 · CVPR 2014 6 · N. Siddharth, Andrei Barbu, Jeffrey Mark Siskind

We present a system that demonstrates how the compositional structure of events, in concert with the compositional structure of language, can interplay with the underlying focusing mechanisms in video action recognition,…

Action RecognitionActivity RecognitionSentenceTemporal Action Localization

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

2025-08-31 · Xiangchen Wang, Jinrui Zhang, Teng Wang, Haigang Zhang 외 arxiv

Recent advancements in large video-language models have revolutionized video understanding tasks. However, their efficiency is significantly constrained by processing high volumes of visual tokens. Existing token compres…

Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert

2023-03-29 · CVPR 2023 1 · Jiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan 외

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and vis…

Contrastive LearningFace GenerationLip ReadingTalking Face Generation