paper-with-me

홈 › Papers

MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose

2026-02-22 · Sirine Bhouri, Lan Wei, Jian-Qing Zheng, Dandan Zhang arxiv

Acquiring aligned visuo-tactile datasets is slow and costly, requiring specialised hardware and large-scale data collection. Synthetic generation is promising, but prior methods are typically single-modality, limiting cross-modal learning. We present MultiDiffSense, a unified diffusion model that synthesises images for multiple vision-based tactile sensors (ViTac, TacTip, ViTacTip) within a single architecture. Our approach uses dual conditioning on CAD-derived, pose-aligned depth maps and structured prompts that encode sensor type and 4-DoF contact pose, enabling controllable, physically consistent multi-modal synthesis. Evaluating on 8 objects (5 seen, 3 novel) and unseen poses, MultiDiffSense outperforms a Pix2Pix cGAN baseline in SSIM by +36.3% (ViTac), +134.6% (ViTacTip), and +64.7% (TacTip). For downstream 3-DoF pose estimation, mixing 50% synthetic with 50% real halves the required real data while maintaining competitive performance. MultiDiffSense alleviates the data-collection bottleneck in tactile sensing and enables scalable, controllable multi-modal dataset generation for robotic applications.

📄 PDF Abstract BibTeX arXiv:2602.19348

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationPose Estimation

Similar Papers 제목 키워드 기반

Multimodal and Force-Matched Imitation Learning with a See-Through Visuotactile Sensor

2023-11-02 · Trevor Ablett, Oliver Limoyo, Adam Sigal, Affan Jilani 외

Contact-rich tasks continue to present many challenges for robotic manipulation. In this work, we leverage a multimodal visuotactile sensor within the framework of imitation learning (IL) to perform contact-rich tasks th…

Imitation LearningSTS

ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

2025-05-26 · Fotios Lygerakis, Ozan Özdenizci, Elmar Rückert

Tactile sensing provides local essential information that is complementary to visual perception, such as texture, compliance, and force. Despite recent advances in visuotactile representation learning, challenges remain …

cross-modal alignmentPositionRepresentation LearningRobotic Grasping+3

Inference-time Policy Steering via Vision and Touch

2026-06-12 · Yilin Wu, Zilin Si, Zeynep Temel, Oliver Kroemer 외 arxiv

Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution. While prior methods typically perform this verification only with visual observation…

Contact-Grounded Policy: Dexterous Visuotactile Policy with Generative Contact Grounding

2026-03-05 · Zhengtong Xu, Yeping Wang, Ben Abbatematteo, Jom Preechayasomboon 외 arxiv

Contact-rich dexterous manipulation with multi-finger hands remains an open challenge in robotics because task success depends on multi-point contacts that continuously evolve and are highly sensitive to object geometry,…

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

2026-01-28 · Zhemeng Zhang, Jiahua Ma, Xincheng Yang, Xin Wen 외 arxiv

Fine-grained and contact-rich manipulation remain challenging for robots, largely due to the underutilization of tactile feedback. To address this, we introduce TouchGuide, a novel cross-policy visuo-tactile fusion parad…

Contrastive Learning