paper-with-me

홈 › Papers

ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning

2026-02-12 · Yufeng Tian, Shuiqi Cheng, Tianming Wei, Tianxing Zhou, Yuanhang Zhang, Zixian Liu, Qianwei Han, Zhecheng Yuan, Huazhe Xu arxiv

Tactile information plays a crucial role in human manipulation tasks and has recently garnered increasing attention in robotic manipulation. However, existing approaches mostly focus on the alignment of visual and tactile features and the integration mechanism tends to be direct concatenation. Consequently, they struggle to effectively cope with occluded scenarios due to neglecting the inherent complementary nature of both modalities and the alignment may not be exploited enough, limiting the potential of their real-world deployment. In this paper, we present ViTaS, a simple yet effective framework that incorporates both visual and tactile information to guide the behavior of an agent. We introduce Soft Fusion Contrastive Learning, an advanced version of conventional contrastive learning method and a CVAE module to utilize the alignment and complementarity within visuo-tactile representations. We demonstrate the effectiveness of our method in 12 simulated and 3 real-world environments, and our experiments show that ViTaS significantly outperforms existing baselines. Project page: https://skyrainwind.github.io/ViTaS/index.html.

📄 PDF Abstract BibTeX arXiv:2602.11643

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing

2026-04-02 · Junyi Zong, Qingxuan Jia, Meixian Shi, Tong Li 외 arxiv

Quality inspection in smart manufacturing requires identifying intrinsic material and surface properties beyond visible geometry, yet vision-only methods remain vulnerable to occlusion and reflection. We propose VitaTouc…

Contrastive LearningSemantic Similarity

ViTaSCOPE: Visuo-tactile Implicit Representation for In-hand Pose and Extrinsic Contact Estimation

2025-06-13 · Jayjun Lee, Nima Fazeli

Mastering dexterous, contact-rich object manipulation demands precise estimation of both in-hand object poses and external contact locations$\unicode{x2013}$tasks particularly challenging due to partial and noisy observa…

3D geometryObjectPose Estimation

ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations

2025-06-25 · Zhiyuan Wu, Yongqiang Zhao, Shan Luo

Vision and touch are two fundamental sensory modalities for robots, offering complementary information that enhances perception and manipulation tasks. Previous research has attempted to jointly learn visual-tactile repr…

Contrastive LearningMaterial ClassificationRepresentation Learning

From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning

2024-10-09 · Yang Bai, Yang Zhou, Jun Zhou, Rick Siow Mong Goh 외

Large vision language models (VLMs) combine large language models with vision encoders, demonstrating promise across various tasks. However, they often underperform in task-specific applications due to domain gaps betwee…

Medical Diagnosis

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

2026-01-28 · Zhemeng Zhang, Jiahua Ma, Xincheng Yang, Xin Wen 외 arxiv

Fine-grained and contact-rich manipulation remain challenging for robots, largely due to the underutilization of tactile feedback. To address this, we introduce TouchGuide, a novel cross-policy visuo-tactile fusion parad…

Contrastive Learning