paper-with-me

홈 › Papers

Learning 4D Geometric Priors for Inference-Efficient World Action Models

2026-07-06 · Jianjun Zhang, Jian Zhu, Taiyi Su, Chong Ma, Zitai Huang, Yi Xu, Hanli Wang arxiv

World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise manipulation. We propose MECo-WAM, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. During training, MECo-WAM combines video and action experts with a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder. Asymmetric expert visibility prevents non-causal shortcuts from auxiliary geometry to action generation. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency. We further propose action-aware temporal geometric distillation, which aligns within-frame geometric relations and their temporal evolution while emphasizing visual regions most relevant to robot actions. At deployment, all auxiliary 4D components are removed. Experiments on LIBERO (98.2%), RoboTwin 2.0 (92.6%), and challenging real-world manipulation tasks show that MECo-WAM improves manipulation performance without increasing inference cost.

📄 PDF Abstract BibTeX arXiv:2607.05468

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

2026-06-12 · Ying Li, Xiaobao Wei, Jiajun Cao, Hao Wang 외 arxiv

World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video or latent spaces, where visually plausibl…

Inference with correlated priors using sisters cells

2025-05-20 · Sina Tootoonian, Andreas T. Schaefer

A common view of sensory processing is as probabilistic inference of latent causes from receptor activations. Standard approaches often assume these causes are a priori independent, yet real-world generative factors are …

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

2026-04-29 · Jun Guo, Qiwei Li, Peiyan Li, Zilong Chen 외 arxiv

We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruction) in a single framework, addressing the critical limitations of pr…

3D Reconstruction

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs

2026-04-07 · Chongyu Wang, Ting Huang, Chunyu Sun, Xinyu Ning 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in 2D visual tasks but still exhibit limited physical spatial awareness when processing real-world visual streams. Recently, feed-forward geometr…

Spatial Reasoning

Geometric Action Model for Robot Policy Learning

2026-06-15 · Jisang Han, Seonghu Jeon, Jaewoo Jung, René Zurbrügg 외 arxiv

Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-action models (VLAs) and video world-action …

Robot Manipulation