paper-with-me

홈 › Papers

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

2026-06-12 · Ying Li, Xiaobao Wei, Jiajun Cao, Hao Wang, Xiaowei Chi, Chengyu Bai, Qianpu Sun, Jiajun Li, Xiaojie Zhang, Peidong Jia, Jian Tang, Sirui Han, Shanghang Zhang arxiv

World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video or latent spaces, where visually plausible rollouts miss the 3D spatial constraints and occluded contact geometry required for precise manipulation. While geometric foundation models offer strong priors for recovering dense 3D structure and motion from visual observations, forcing WAMs to predict the dense 4D representation introduces costly geometric decoding and slows down causal action generation. To address the trade-off, we present WAM4D, a fast 4D world action model that uses lightweight spatial register tokens as training-time future-depth readouts to transfer pretrained geometric priors into a causal video-action transformer, then removes the register branch for lightweight action inference. To prevent non-causal shortcuts, we further design causal mixture attention for the Mixture-of-Transformers (MoT) WAM backbone, defining modality-specific visibility among video, action, and geometry tokens. Comprehensive experiments on RoboTwin 2.0 and challenging real-world manipulation tasks show that WAM4D improves spatial consistency and achieves competitive action prediction while maintaining efficient inference.

📄 PDF Abstract BibTeX arXiv:2606.14048

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models

2025-09-25 · Jiyeon Koo, Taewan Cho, Hyunjoon Kang, Eunseom Pyo 외 arxiv

Vision-Language-Action (VLA) models have demonstrated robust performance across diverse robotic tasks. However, their high memory and computational demands often limit real-time deployment. While existing model compressi…

Spatial ReasoningModel Compression

Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

2024-10-17 · Yuxin Wen, Qingqing Cao, Qichen Fu, Sachin Mehta 외

Recent advancements in vision-language models (VLMs) have expanded their potential for real-world applications, enabling these models to perform complex reasoning on images. In the widely used fully autoregressive transf…

Computational Efficiency

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

2026-05-19 · Congpei Qiu, Zhaoyu Hu, Wei Ke, Zhuotao Tian 외 arxiv

Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hindered by spurious tokens. Prior efforts to mitigate this have been li…

Representation Learning

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

2025-01-27 · Renshan Zhang, Rui Shao, Gongwei Chen, Kaiwen Zhou 외

The incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a …

Convolutions Need Registers Too: HVS-Inspired Dynamic Attention for Video Quality Assessment

2026-01-16 · Mayesha Maliha R. Mithila, Mylene C. Q. Farias arxiv

No-reference video quality assessment (NR-VQA) estimates perceptual quality without a reference video, which is often challenging. While recent techniques leverage saliency or transformer attention, they merely address g…

Video Quality AssessmentSaliency Prediction