paper-with-me

Papers

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction

2026-06-05 · Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung, Jaejin Lee, Taesup Kim arxiv

Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces. We study this question through a unified probe-based evaluation across diverse encoder families, including image-only self-supervision, video pretraining with and without latent prediction, reconstruction-based autoencoders, diffusion models, and shortcut-forcing dynamics models. Using a common inverse-dynamics probing objective, we find that action-relevant structure is driven primarily by temporal video pretraining rather than pixel reconstruction fidelity: models with strong pixel decoding quality can exhibit near-zero action recoverability, while video-pretrained self-supervised encoders consistently achieve the best Pareto trade-off between visual fidelity and action prediction. Comparing V-JEPA and VideoMAE further shows that most gains arise from natural-video temporal context, with feature-level latent prediction providing a smaller additional benefit. These trends transfer across robotic benchmarks, though CALVIN reveals that static-environment tasks can partially mask the importance of temporal structure by allowing strong image priors to suffice. Finally, inverse-dynamics supervision substantially improves robustness to visual corruption, suggesting that action-aware objectives regularize latent geometry beyond clean-setting performance. Our results identify temporal predictive structure -- not reconstruction fidelity -- as the primary ingredient underlying action-relevant video representations.

📄 PDF Abstract BibTeX arXiv:2606.07687

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

2026-06-08 · Jiajun Li, Tiecheng Guo, Yifan Ye, Rongyu Zhang 외 arxiv

World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, whic…

Video Prediction

SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation

2026-03-17 · Jiongze Yu, Xiangbo Gao, Pooja Verlani, Akshay Gadde 외 arxiv

Video Super-Resolution (VSR) aims to restore high-quality video frames from low-resolution (LR) estimates, yet most existing VSR approaches behave like black boxes at inference time: users cannot reliably correct unexpec…

Image Super-ResolutionVideo Super-ResolutionStyle Transfer

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

2026-08-06 · Zhe Li, Zhenzhe Zhang, Yangyang Wei, Wenjie Zhang 외 arxiv

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies ty…

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

2026-06-09 · Xiaoquan Sun, Ruijian Zhang, Chen Cao, Yihan Sun 외 arxiv

World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and robustness. However, existing WAMs stil…

Optical Flow EstimationCausal Inference

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

2026-08-21 · Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang 외 arxiv

Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion a…

Autonomous Driving