paper-with-me

홈 › Papers

TLA: Tactile-Language-Action Model for Contact-Rich Manipulation

2025-03-11 · Peng Hao, Chaofan Zhang, Dingzhe Li, Xiaoge Cao, Xiaoshuai Hao, Shaowei Cui, Shuo Wang

Significant progress has been made in vision-language models. However, language-conditioned robotic manipulation for contact-rich tasks remains underexplored, particularly in terms of tactile sensing. To address this gap, we introduce the Tactile-Language-Action (TLA) model, which effectively processes sequential tactile feedback via cross-modal language grounding to enable robust policy generation in contact-intensive scenarios. In addition, we construct a comprehensive dataset that contains 24k pairs of tactile action instruction data, customized for fingertip peg-in-hole assembly, providing essential resources for TLA training and evaluation. Our results show that TLA significantly outperforms traditional imitation learning methods (e.g., diffusion policy) in terms of effective action generation and action accuracy, while demonstrating strong generalization capabilities by achieving over 85\% success rate on previously unseen assembly clearances and peg shapes. We publicly release all data and code in the hope of advancing research in language-conditioned tactile manipulation skill learning. Project website: https://sites.google.com/view/tactile-language-action/

📄 PDF Abstract BibTeX arXiv:2503.08548

Code (0)

등록된 구현이 없습니다.

Tasks

Action GenerationContact-rich ManipulationImitation Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
TLA Please enter a description about the method here

Similar Papers 제목 키워드 기반

UniTacVLA: Unified Tactile Understanding and Prediction in Vision Language Action Models

2026-06-30 · Xidong Zhang, Yichi Zhang, Jiaxin Shi, Fucai Zhu 외 arxiv

Vision-language-action (VLA) models have achieved strong performance in many robotic manipulation tasks, yet remain limited in contact-rich dexterous manipulation. To overcome this limitation, recent vision-tactile-langu…

ReTouch: Empowering Contact-Rich Dexterous Manipulation with Online-Refined Tactile Prediction

2026-08-03 · Shiqi Zhang, Xin Zhang, Yedong Shen, Yao Li 외 arxiv

Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effectively integrating tactile feedback into…

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

2026-07-26 · NeoteAI Team, Fudan TEAI Team hf

We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large sc…

Video Prediction

Representation-Aligned Tactile Grounding for Contact-Rich Robotic Manipulation

2026-07-16 · Ruilin Chen, Jingkai Jia, Tong Yang, Xinyu Zhou 외 arxiv

Tactile-enhanced vision-language-action (VLA) policies have been introduced for contact-rich manipulation, where critical interaction states are often hidden from vision. Future tactile prediction is a promising way to u…

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

2026-07-02 · Shuai Tian, Yupeng Zheng, Yuhang Zheng, Songen Gu 외 arxiv

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observations. Existing visual-tactile policies u…