paper-with-me

Papers

TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception

2025-11-24 · Kailin Lyu, Long Xiao, Jianing Zeng, Junhao Dong, Xuexin Liu, Zhuojun Zou, Haoyue Yang, Lin Shu, Jie Hao arxiv

Traditional vision-based material perception methods often experience substantial performance degradation under visually impaired conditions, thereby motivating the shift toward non-visual multimodal material perception. Despite this, existing approaches frequently perform naive fusion of multimodal inputs, overlooking key challenges such as modality-specific noise, missing modalities common in real-world scenarios, and the dynamically varying importance of each modality depending on the task. These limitations lead to suboptimal performance across several benchmark tasks. In this paper, we propose a robust multimodal fusion framework, TouchFormer. Specifically, we employ a Modality-Adaptive Gating (MAG) mechanism and intra- and inter-modality attention mechanisms to adaptively integrate cross-modal features, enhancing model robustness. Additionally, we introduce a Cross-Instance Embedding Regularization(CER) strategy, which significantly improves classification accuracy in fine-grained subcategory material recognition tasks. Experimental results demonstrate that, compared to existing non-visual methods, the proposed TouchFormer framework achieves classification accuracy improvements of 2.48% and 6.83% on SSMC and USMC tasks, respectively. Furthermore, real-world robotic experiments validate TouchFormer's effectiveness in enabling robots to better perceive and interpret their environment, paving the way for its deployment in safety-critical applications such as emergency response and industrial automation. The code and datasets will be open-source, and the videos are available in the supplementary materials.

📄 PDF Abstract BibTeX arXiv:2511.19509

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Meta-Transformer: A Unified Framework for Multimodal Learning

2023-07-20 · Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li 외

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging to design a unified network for processi…

Time Series

Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features

2025-08-07 · Manish Kansana, Elias Hossain, Shahram Rahimi, Noorbakhsh Amiri Golilarz arxiv

Surface material recognition is a key component in robotic perception and physical interaction, particularly when leveraging both tactile and visual sensory inputs. In this work, we propose Surformer v1, a transformer-ba…

Computational EfficiencyFeature Engineering

Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision

2025-09-04 · Manish Kansana, Sindhuja Penchala, Shahram Rahimi, Noorbakhsh Amiri Golilarz arxiv

Multimodal surface material classification plays a critical role in advancing tactile perception for robotic manipulation and interaction. In this paper, we present Surformer v2, an enhanced multi-modal classification ar…

Multi-modal Classification

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

2025-10-13 · Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino arxiv

We introduce J-ORA, a novel multimodal dataset that bridges the gap in robot perception by providing detailed object attribute annotations within Japanese human-robot dialogue scenarios. J-ORA is designed to support thre…

Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation

2025-12-10 · Yuyang Li, Yinghan Chen, Zihang Zhao, Puhao Li 외 arxiv

Robotic manipulation requires both rich multimodal perception and effective learning frameworks to handle complex real-world tasks. See-through-skin (STS) sensors, which combine tactile and visual perception, offer promi…

Robot ManipulationContact Detection