paper-with-me

홈 › Papers

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

2026-03-24 · Haoran Yuan, Weigang Yi, Zhenyu Zhang, Wendi Chen, Yuchen Mo, Jiashi Yin, Xinzhuo Li, Xiangyu Zeng, Chuan Wen, Cewu Lu, Katherine Driggs-Campbell, Ismini Lourentzou arxiv

Video-Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent action predictions. Although such models demonstrate strong performance on long-horizon tasks through visual reasoning, they remain limited in contact-rich scenarios where critical interaction states are only partially observable from vision alone. In particular, fine-grained force modulation and contact transitions are not reliably encoded in visual tokens, leading to unstable or imprecise behaviors. To bridge this gap, we introduce the Video-Tactile Action Model (VTAM), a multimodal world modeling framework that incorporates tactile perception as a complementary grounding signal. VTAM augments a pretrained video transformer with tactile streams via a lightweight modality transfer finetuning, enabling efficient cross-modal representation learning without tactile-language paired data or independent tactile pretraining. To stabilize multimodal fusion, we introduce a tactile regularization loss that enforces balanced cross-modal attention, preventing visual latent dominance in the action model. VTAM demonstrates superior performance in contact-rich manipulation, maintaining a robust success rate of 90 percent on average. In challenging scenarios such as potato chip pick-and-place requiring high-fidelity force awareness, VTAM outperforms the pi 0.5 baseline by 80 percent. Our findings demonstrate that integrating tactile feedback is essential for correcting visual estimation errors in world action models, providing a scalable approach to physically grounded embodied foundation models.

📄 PDF Abstract BibTeX arXiv:2603.23481

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningVisual Reasoning

Similar Papers 제목 키워드 기반

MVTamperBench: Evaluating Robustness of Vision-Language Models

2024-12-27 · Amit Agarwal, Srikant Panda, Angeline Charles, Bhargava Kumar 외

Multimodal Large Language Models (MLLMs) have driven major advances in video understanding, yet their vulnerability to adversarial tampering and manipulations remains underexplored. To address this gap, we introduce MVTa…

Video Understanding

Combining Vision and Tactile Sensation for Video Prediction

2023-04-21 · Willow Mandil, Amir Ghalamzan-E

In this paper, we explore the impact of adding tactile sensation to video prediction models for physical robot interactions. Predicting the impact of robotic actions on the environment is a fundamental challenge in robot…

PredictionVideo Prediction

Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction

2026-09-07 · Ruoqu Chen, Feixiang Ruan, Liu Cao, Zihao Wang 외 arxiv

Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable vis…

Zero-shot Generalization

Universal Visuo-Tactile Video Understanding for Embodied Interaction

2025-05-28 · Yifan Xie, Mingyang Li, Shoujie Li, Xingting Li 외

Tactile perception is essential for embodied agents to understand physical attributes of objects that cannot be determined through visual inspection alone. While existing approaches have made progress in visual and langu…

FrictionLarge Language ModelText GenerationVideo Understanding

Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention

2026-06-25 · Siyu Wu, Linjing You, Junjie Zhu, Yaozu Liu 외 arxiv

World Action Models (WAMs) jointly predict future visual observations and actions, but visual futures alone often miss slip, jamming, contact-direction changes, and subtle misalign- ment in contact-rich manipulation. Tac…

Decision Making