paper-with-me

Papers

Visual-Tactile Cross-Modal Data Generation using Residue-Fusion GAN with Feature-Matching and Perceptual Losses

2021-07-12 · Shaoyu Cai, Kening Zhu, Yuki Ban, Takuji Narumi

Existing psychophysical studies have revealed that the cross-modal visual-tactile perception is common for humans performing daily activities. However, it is still challenging to build the algorithmic mapping from one modality space to another, namely the cross-modal visual-tactile data translation/generation, which could be potentially important for robotic operation. In this paper, we propose a deep-learning-based approach for cross-modal visual-tactile data generation by leveraging the framework of the generative adversarial networks (GANs). Our approach takes the visual image of a material surface as the visual data, and the accelerometer signal induced by the pen-sliding movement on the surface as the tactile data. We adopt the conditional-GAN (cGAN) structure together with the residue-fusion (RF) module, and train the model with the additional feature-matching (FM) and perceptual losses to achieve the cross-modal data generation. The experimental results show that the inclusion of the RF module, and the FM and the perceptual losses significantly improves cross-modal data generation performance in terms of the classification accuracy upon the generated data and the visual similarity between the ground-truth and the generated data.

📄 PDF Abstract BibTeX arXiv:2107.05468

Code (1)

shaoyuca/Visual-Tactile-Data-Generation 공식 구현 tf

Tasks

Translation

Similar Papers 제목 키워드 기반

Tactile DreamFusion: Exploiting Tactile Sensing for 3D Generation

2024-12-09 · Ruihan Gao, Kangle Deng, Gengshan Yang, Wenzhen Yuan 외

3D generation methods have shown visually compelling results powered by diffusion image priors. However, they often fail to produce realistic geometric details, resulting in overly smooth surfaces or geometric details in…

3D GenerationImage to 3DText to 3DTexture Synthesis

Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models

2026-08-01 · Zihang Yao, Chaoyue Ding, Yingying Yu arxiv

Contact-rich manipulation remains challenging because successful control depends on physical interaction cues that are often weakly observable from vision alone. Recent tactile world action models jointly model future vi…

Universal Visuo-Tactile Video Understanding for Embodied Interaction

2025-05-28 · Yifan Xie, Mingyang Li, Shoujie Li, Xingting Li 외

Tactile perception is essential for embodied agents to understand physical attributes of objects that cannot be determined through visual inspection alone. While existing approaches have made progress in visual and langu…

FrictionLarge Language ModelText GenerationVideo Understanding

OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning

2026-01-01 · Liuxiang Qiu, Hui Da, Yuzhen Niu, Tiesong Zhao 외 arxiv

Visual-tactile learning (VTL) enables embodied agents to perceive the physical world by integrating visual (VIS) and tactile (TAC) sensors. However, VTL still suffers from modality discrepancies between VIS and TAC image…

Domain Generalization

TextToucher: Fine-Grained Text-to-Touch Generation

2024-09-09 · Jiahang Tu, Hao Fu, Fengyu Yang, Hanbin Zhao 외

Tactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tacti…

Language ModellingLarge Language ModelMultimodal Large Language Model