paper-with-me

Papers

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

2026-07-16 · Wei Li, Peijin Jia, Yuan Ma, Xuefeng Jiang, Titong Jiang, Sheng Sun, Yujian Li, Xin Wen, Han Hong, Zhikang Liu, Bailin Li, Kun Zhan arxiv

Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue that future feature prediction and sparse point tracking are naturally complementary: the former provides the goal state, while the latter captures the continuous motion path toward it. We propose FoMoVLA, a framework that augments VLA representations with explicit spatio-temporal supervision by jointly learning future feature foresight and sparse 2D point tracking, enhancing the continuous action policy. FoMoVLA introduces compact foresight tokens to decode future feature states, decodes sparse temporal 2D point trajectories to model compact geometric motion, and couples both through a lightweight future-conditioned cross-attention module that enables consistent reasoning between anticipated states and point dynamics. Extensive experiments on LIBERO, RoboCasa GR-1 Tabletop, and LIBERO-Plus demonstrate state-of-the-art performance and strong zero-shot generalization. Project page is available at https://liauto-research.github.io/FoMoVLA.

📄 PDF Abstract BibTeX arXiv:2607.14739

Code (3)

BaiShuanghao/my_arXiv_daily ★ 201
InsomaniacElf/sg-tamil-tts-resources- ★ 1
Tavish9/awesome-daily-AI-arxiv ★ 111

Tasks

Zero-shot GeneralizationPoint Tracking

Similar Papers 제목 키워드 기반

ForeSightGuide: An Anticipatory Framework toward Accurate and Low-Redundancy Guidance for the Visually Impaired

2026-08-19 · Zhiyuan Wang, Xu Li, Shikang Guo, Wei Meng 외 arxiv

Electronic travel aids are pivotal for the independent mobility of the visually impaired. While Vision-Language Models (VLMs) offer rich environmental understanding, they often suffer from excessive false positives in dy…

Scene Understanding

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

2026-08-05 · Zehua Fan, Junjie He, Wenxuan Song, Xi Wang 외 arxiv

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body mani…

Video Generation

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

2025-09-08 · Qi Lv, Weijie Kong, Hao Li, Jia Zeng 외 arxiv

Executing language-conditioned tasks in dynamic visual environments remains a central challenge in embodied AI. Existing Vision-Language-Action (VLA) models predominantly adopt reactive state-to-action mappings, often le…

Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation

2026-03-13 · Minghao Jin, Mozheng Liao, Mingfei Han, Zhihui Li 외 arxiv

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates er…

ForeAct: Steering Your VLA with Efficient Visual Foresight Planning

2026-02-12 · Zhuoyang Zhang, Shang Yang, Qinghao Hu, Luke J. Huang 외 arxiv

Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We present Visual Foresight Planning (Fore…

Image Generation