paper-with-me

Papers

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

2024-01-31 · CVPR 2024 1 · Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, Alex Wong

The ability to associate touch with other modalities has huge implications for humans and computational systems. However, multimodal learning with touch remains challenging due to the expensive data collection process and non-standardized sensor outputs. We introduce UniTouch, a unified tactile model for vision-based touch sensors connected to multiple modalities, including vision, language, and sound. We achieve this by aligning our UniTouch embeddings to pretrained image embeddings already associated with a variety of other modalities. We further propose learnable sensor-specific tokens, allowing the model to learn from a set of heterogeneous tactile sensors, all at the same time. UniTouch is capable of conducting various touch sensing tasks in the zero-shot setting, from robot grasping prediction to touch image question answering. To the best of our knowledge, UniTouch is the first to demonstrate such capabilities. Project page: https://cfeng16.github.io/UniTouch/

📄 PDF Abstract BibTeX arXiv:2401.18084

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Touch-R1: Reinforcing Touch Reasoning in MLLMs

2026-05-26 · Yingxin Lai, Yafei Zhou, Fucai Zhu, Siyu Zhu 외 arxiv

While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile-language models primarily rely on supervised or co…

Reinforcement Learning

VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios

2026-07-16 · Kailin Lyu, Long Xiao, Jianing Zeng, Di Wu 외 arxiv

Tactile image generation significantly reduces the dependency on expensive and wear-prone sensors by synthesizing high-fidelity tactile data, offering an efficient solution for tactile information acquisition in robotic …

multimodal generationImage Generation

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

2026-06-30 · Jiahang Tu, Fengyu Yang, Chenyang Ma, Xihang Yu 외 arxiv

Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, existing research rarely extends this paradigm to the tactile domain, where both o…

A Touch, Vision, and Language Dataset for Multimodal Alignment

2024-02-20 · Letian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch 외

Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactil…

Language ModelingLanguage ModellingText Generation

TextToucher: Fine-Grained Text-to-Touch Generation

2024-09-09 · Jiahang Tu, Hao Fu, Fengyu Yang, Hanbin Zhao 외

Tactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tacti…

Language ModellingLarge Language ModelMultimodal Large Language Model