paper-with-me

홈 › Papers

A Touch, Vision, and Language Dataset for Multimodal Alignment

2024-02-20 · Letian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch, Jaimyn Drake, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, Ken Goldberg

Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the complexity of aligning tactile readings with both visual observations and language descriptions. As a step towards bridging that gap, this work introduces a new dataset of 44K in-the-wild vision-touch pairs, with English language labels annotated by humans (10%) and textual pseudo-labels from GPT-4V (90%). We use this dataset to train a vision-language-aligned tactile encoder for open-vocabulary classification and a touch-vision-language (TVL) model for text generation using the trained encoder. Results suggest that by incorporating touch, the TVL model improves (+29% classification accuracy) touch-vision-language alignment over existing models trained on any pair of those modalities. Although only a small fraction of the dataset is human-labeled, the TVL model demonstrates improved visual-tactile understanding over GPT-4V (+12%) and open-source vision-language models (+32%) on a new touch-vision understanding benchmark. Code and data: https://tactile-vlm.github.io.

📄 PDF Abstract BibTeX arXiv:2402.13232

Code (1)

Max-Fu/tvl 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingText Generation

Similar Papers 제목 키워드 기반

Towards Comprehensive Multimodal Perception: Introducing the Touch-Language-Vision Dataset

2024-03-14 · Ning Cheng, You Li, Jing Gao, Bin Fang 외

Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile…

Sentence

OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction

2025-12-18 · Yuxin Ray Song, Jinzhou Li, Rao Fu, Devin Murphy 외 arxiv

The human hand is our primary interface to the physical world, yet egocentric perception rarely knows when, where, or how forcefully it makes contact. Robust wearable tactile sensors are scarce, and no existing in-the-wi…

TactEx: An Explainable Multimodal Robotic Interaction Framework for Human-Like Touch and Hardness Estimation

2026-02-21 · Felix Verstraete, Lan Wei, Wen Fan, Dandan Zhang arxiv

Accurate perception of object hardness is essential for safe and dexterous contact-rich robotic manipulation. Here, we present TactEx, an explainable multimodal robotic interaction framework that unifies vision, touch, a…

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

2024-01-31 · CVPR 2024 1 · Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park 외

The ability to associate touch with other modalities has huge implications for humans and computational systems. However, multimodal learning with touch remains challenging due to the expensive data collection process an…

Question AnsweringVisual Question Answering (VQA)

VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing

2026-04-02 · Junyi Zong, Qingxuan Jia, Meixian Shi, Tong Li 외 arxiv

Quality inspection in smart manufacturing requires identifying intrinsic material and surface properties beyond visible geometry, yet vision-only methods remain vulnerable to occlusion and reflection. We propose VitaTouc…

Contrastive LearningSemantic Similarity