paper-with-me

홈 › Papers

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

2026-07-31 · Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li hf

World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.

📄 PDF Abstract BibTeX arXiv:2607.28993

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation

2026-03-14 · You Wu, Zixuan Chen, Cunxu Ou, Wenxuan Wang 외 arxiv

Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA) frameworks typically use 2D representati…

Continuous ControlRobot Manipulation

Semantic Decomposition and Recognition of Long and Complex Manipulation Action Sequences

2016-10-18 · Eren Erdal Aksoy, Adil Orhan, Florentin Woergoetter

Understanding continuous human actions is a non-trivial but important problem in computer vision. Although there exists a large corpus of work in the recognition of action sequences, most approaches suffer from problems …

Semantic Segmentation

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation

2026-04-29 · Yuxuan Tian, Yurun Jin, Bin Yu, Yukun Shi 외 arxiv

Robotic manipulation requires reasoning about future spatial-temporal interactions and geometric constraints, yet existing Vision-Language-Action (VLA) policies often leave predictive representation weakly coupled with a…

HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data

2025-10-13 · Ruizhe Liu, Pei Zhou, Qian Luo, Li Sun 외 arxiv

Effective generalization in robotic manipulation requires representations that capture invariant patterns of interaction across environments and tasks. We present a self-supervised framework for learning hierarchical man…

Representation Learning

Geometric Action Model for Robot Policy Learning

2026-06-15 · Jisang Han, Seonghu Jeon, Jaewoo Jung, René Zurbrügg 외 arxiv

Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-action models (VLAs) and video world-action …

Robot Manipulation