paper-with-me

Papers

Towards Comprehensive Multimodal Perception: Introducing the Touch-Language-Vision Dataset

2024-03-14 · Ning Cheng, You Li, Jing Gao, Bin Fang, Jinan Xu, Wenjuan Han

Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile modalities, with limited exploration in the domain of language. Beyond vocabulary, sentence-level descriptions contain richer semantics. Based on this, we construct a touch-language-vision dataset named TLV (Touch-Language-Vision) by human-machine cascade collaboration, featuring sentence-level descriptions for multimode alignment. The new dataset is used to fine-tune our proposed lightweight training framework, STLV-Align (Synergistic Touch-Language-Vision Alignment), achieving effective semantic alignment with minimal parameter adjustments (1%). Project Page: https://xiaoen0.github.io/touch.page/.

📄 PDF Abstract BibTeX arXiv:2403.09813

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Multimodal perception for dexterous manipulation

2021-12-28 · Guanqun Cao, Shan Luo

Humans usually perceive the world in a multimodal way that vision, touch, sound are utilised to understand surroundings from various dimensions. These senses are combined together to achieve a synergistic effect where th…

3D ReconstructionFrictionTranslation

TactEx: An Explainable Multimodal Robotic Interaction Framework for Human-Like Touch and Hardness Estimation

2026-02-21 · Felix Verstraete, Lan Wei, Wen Fan, Dandan Zhang arxiv

Accurate perception of object hardness is essential for safe and dexterous contact-rich robotic manipulation. Here, we present TactEx, an explainable multimodal robotic interaction framework that unifies vision, touch, a…

OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction

2025-12-18 · Yuxin Ray Song, Jinzhou Li, Rao Fu, Devin Murphy 외 arxiv

The human hand is our primary interface to the physical world, yet egocentric perception rarely knows when, where, or how forcefully it makes contact. Robust wearable tactile sensors are scarce, and no existing in-the-wi…

TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception

2025-11-24 · Kailin Lyu, Long Xiao, Jianing Zeng, Junhao Dong 외 arxiv

Traditional vision-based material perception methods often experience substantial performance degradation under visually impaired conditions, thereby motivating the shift toward non-visual multimodal material perception.…

Touch-R1: Reinforcing Touch Reasoning in MLLMs

2026-05-26 · Yingxin Lai, Yafei Zhou, Fucai Zhu, Siyu Zhu 외 arxiv

While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile-language models primarily rely on supervised or co…

Reinforcement Learning