paper-with-me

Papers

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation

2025-08-30 · Chuye Zhang, Xiaoxiong Zhang, Wei Pan, Linfang Zheng, Wei Zhang arxiv

Robotic manipulation in unstructured environments requires systems that can generalize across diverse tasks while maintaining robust and reliable performance. We introduce {GVF-TAPE}, a closed-loop framework that combines generative visual foresight with task-agnostic pose estimation to enable scalable robotic manipulation. GVF-TAPE employs a generative video model to predict future RGB-D frames from a single side-view RGB image and a task description, offering visual plans that guide robot actions. A decoupled pose estimation model then extracts end-effector poses from the predicted frames, translating them into executable commands via low-level controllers. By iteratively integrating video foresight and pose estimation in a closed loop, GVF-TAPE achieves real-time, adaptive manipulation across a broad range of tasks. Extensive experiments in both simulation and real-world settings demonstrate that our approach reduces reliance on task-specific action data and generalizes effectively, providing a practical and scalable solution for intelligent robotic systems.

📄 PDF Abstract BibTeX arXiv:2509.00361

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

Mirai: Autoregressive Visual Generation Needs Foresight

2026-01-21 · Yonghao Yu, Lang Huang, Zerun Wang, Runyi Li 외 arxiv

Autoregressive (AR) visual generators model images as sequences of discrete tokens and are trained with a next-token likelihood objective. This strict causal supervision optimizes each step based only on the immediate ne…

Image Generation

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

2026-06-25 · Mingyang Lyu, Yinqian Sun, Yiyang Jia, Sicheng Shen 외 arxiv

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward general-purpose task capability, yet their em…

Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation

2026-03-13 · Minghao Jin, Mozheng Liao, Mingfei Han, Zhihui Li 외 arxiv

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates er…

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

2025-09-08 · Qi Lv, Weijie Kong, Hao Li, Jia Zeng 외 arxiv

Executing language-conditioned tasks in dynamic visual environments remains a central challenge in embodied AI. Existing Vision-Language-Action (VLA) models predominantly adopt reactive state-to-action mappings, often le…

ForeAct: Steering Your VLA with Efficient Visual Foresight Planning

2026-02-12 · Zhuoyang Zhang, Shang Yang, Qinghao Hu, Luke J. Huang 외 arxiv

Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We present Visual Foresight Planning (Fore…

Image Generation