paper-with-me

홈 › Papers

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

2026-08-22 · Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren arxiv

World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed. The intermediate frames only describe the visual transition between physical states, which consumes substantial model capacity and computation, but do not directly specify the physical outcome that the robot action is intended to produce. In this paper, we propose DELE-w0.5, which infers robot actions from predicted future states without relying on video generation. Concretely, DELE-w0.5 infers the action sequence from its corresponding compact future latent state. The future latent state captures the action-relevant physical outcome of robot interaction and serves as an explicit bridge between world modeling and action generation. The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame. This formulation removes the high-dimensional visual redundancy introduced by dense video representations, and it therefore enables cheaper training and low-latency inference. Across 640 real-robot trials on four long-horizon manipulation tasks, our DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5% overall full-task success and 81.3% macro ordered-stage progress. It outperforms the strongest baseline by 32.5 percentage points in full-task success and 20.1 percentage points in macro progress.

📄 PDF Abstract BibTeX arXiv:2608.22067

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent Forecasting

2021-03-25 · ICCV 2021 10 · Ye Yuan, Xinshuo Weng, Yanglan Ou, Kris Kitani

Predicting accurate future trajectories of multiple agents is essential for autonomous systems, but is challenging due to the complex agent interaction and the uncertainty in each agent's future behavior. Forecasting mul…

Autonomous DrivingPedestrian Trajectory PredictionTrajectory ForecastingTrajectory Prediction

A Probabilistic Semi-Supervised Approach to Multi-Task Human Activity Modeling

2018-09-24 · Judith Bütepage, Hedvig Kjellström, Danica Kragic

Human behavior is a continuous stochastic spatio-temporal process which is governed by semantic actions and affordances as well as latent factors. Therefore, video-based human activity modeling is concerned with a number…

Action ClassificationGeneral Classificationmotion predictionTrajectory Prediction

The Value of Inferring the Internal State of Traffic Participants for Autonomous Freeway Driving

2017-02-02 · Zachary Sunberg, Christopher Ho, Mykel Kochenderfer

Safe interaction with human drivers is one of the primary challenges for autonomous vehicles. In order to plan driving maneuvers effectively, the vehicle's control system must infer and predict how humans will behave bas…

Autonomous Vehicles

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

2026-08-02 · Ruiteng Zhao, Zhengshen Zhang, Yue Su, Wenshuo Wang 외 hf

World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficie…

On latent position inference from doubly stochastic messaging activities

2012-05-26 · Nam H. Lee, Jordan Yoder, Minh Tang, Carey E. Priebe

We model messaging activities as a hierarchical doubly stochastic point process with three main levels, and develop an iterative algorithm for inferring actors' relative latent positions from a stream of messaging activi…

Position