paper-with-me

홈 › Papers

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

2026-07-02 · Shuai Tian, Yupeng Zheng, Yuhang Zheng, Songen Gu, Yujie Zang, Yuxing Qin, Weize Li, Haoran Li, Wenchao Ding, Dongbin Zhao arxiv

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observations. Existing visual-tactile policies usually feed tactile observations directly into action prediction, but rarely model tactile deformation dynamics during action generation. In this paper, we introduce VT-WAM, a Visual-Tactile World Action Model that jointly learns future visual prediction, tactile deformation prediction, and action prediction within a unified flow matching framework. In particular, VT-WAM introduces (1) Asymmetric Mixture-of-Transformers (MoT) attention to bridge a first-frame visual anchor with temporal tactile dynamics, and (2) contact-gated Action-Visual-Tactile Attention Guidance (AVTAG) to encourage action queries to rely on tactile evidence during contact phases. Across six real-world contact-rich manipulation tasks, VT-WAM achieves a 71.67% average success rate, outperforming Fast-WAM by 26.67% and OmniVTLA by 35.84%. Ablations demonstrate that modeling tactile deformation dynamics and guiding contact-phase tactile attention are both important for contact-rich tasks. Project website: https://vt-wam.github.io/.

📄 PDF Abstract BibTeX arXiv:2607.02503

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

2026-07-04 · Kelin Yu, Haode Zhang, Harish Ravichandar, Yunhai Han 외 hf

Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact…

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

2026-06-07 · Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen 외 arxiv

World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich man…

Robot Manipulation

Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention

2026-06-25 · Siyu Wu, Linjing You, Junjie Zhu, Yaozu Liu 외 arxiv

World Action Models (WAMs) jointly predict future visual observations and actions, but visual futures alone often miss slip, jamming, contact-direction changes, and subtle misalign- ment in contact-rich manipulation. Tac…

Decision Making

TacCoRL: Integrating Tactile Feedback into VLA via Simulation

2026-06-10 · Siyu Ma, Yuqi Liang, Chang Yu, Yunuo Chen 외 arxiv

Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks. We present …

Reinforcement LearningRobot Manipulation

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training

2026-07-03 · Shengbang Liu, Yueru Jia, Yuyang Yan, Jiaming Liu 외 arxiv

Vision-Language-Action (VLA) models have shown promising generalization in robotic manipulation, but they still struggle with contact-rich tasks, where minor contact perturbations can cause unrecoverable failures that ar…