paper-with-me

홈 › Papers

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

2026-06-07 · Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen, Xiaowei Chi, Yaoxu Lyu, Peidong Jia, Sirui Han, Zhihe Lu, Shanghang Zhang arxiv

World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics. Specifically, Dream-Tac introduces (i) contact-gated visuotactile fusion to selectively integrate tactile signals and (ii) a contact-aware attention bias to better regulate cross-modal interactions during manipulation. To support real-time deployment, we further design a dual-level acceleration strategy, reformulating the contact-aware bias to preserve the fused attention path during training and introducing cache-based diffusion acceleration at inference, achieving up to 2.9$\times$ faster training and 1.8$\times$ faster inference. Across six contact-rich manipulation tasks, Dream-Tac improves action accuracy by 31.7\% on average, demonstrating the effectiveness of unified visuotactile world modeling.Code is available at https://github.com/LYFCLOUDFAN/Dream-Tac.

📄 PDF Abstract BibTeX arXiv:2606.08737

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

2025-12-29 · Guo Ye, Zexi Zhang, Xu Zhao, Shang Wu 외 arxiv

Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain blind to physical contact. Consequently, they struggle with contact-rich manipul…

Learning Versatile Humanoid Manipulation with Touch Dreaming

2026-04-14 · Yaru Niu, Zhenlong Fang, Binghong Chen, Shuai Zhou 외 arxiv

Humanoid robots promise general-purpose assistance, yet real-world humanoid loco-manipulation remains challenging because it requires whole-body stability, end-effector dexterity, and contact-aware interaction under freq…

UniTacVLA: Unified Tactile Understanding and Prediction in Vision Language Action Models

2026-06-30 · Xidong Zhang, Yichi Zhang, Jiaxin Shi, Fucai Zhu 외 arxiv

Vision-language-action (VLA) models have achieved strong performance in many robotic manipulation tasks, yet remain limited in contact-rich dexterous manipulation. To overcome this limitation, recent vision-tactile-langu…

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

2026-07-02 · Shuai Tian, Yupeng Zheng, Yuhang Zheng, Songen Gu 외 arxiv

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observations. Existing visual-tactile policies u…

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

2026-07-26 · NeoteAI Team, Fudan TEAI Team hf

We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large sc…

Video Prediction