paper-with-me

Papers

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

2025-02-15 · Ruoxuan Feng, Jiangyu Hu, Wenke Xia, Tianci Gao, Ao Shen, Yuhao Sun, Bin Fang, Di Hu

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding in completing various tasks. However, the distinct data characteristics of these low-standardized visuo-tactile sensors hinder the establishment of a powerful tactile perception system. We consider that the key to addressing this issue lies in learning unified multi-sensor representations, thereby integrating the sensors and promoting tactile knowledge transfer between them. To achieve unified representation of this nature, we introduce TacQuad, an aligned multi-modal multi-sensor tactile dataset from four different visuo-tactile sensors, which enables the explicit integration of various sensors. Recognizing that humans perceive the physical environment by acquiring diverse tactile information such as texture and pressure changes, we further propose to learn unified multi-sensor representations from both static and dynamic perspectives. By integrating tactile images and videos, we present AnyTouch, a unified static-dynamic multi-sensor representation learning framework with a multi-level structure, aimed at both enhancing comprehensive perceptual abilities and enabling effective cross-sensor transfer. This multi-level architecture captures pixel-level details from tactile data via masked modeling and enhances perception and transferability by learning semantic-level sensor-agnostic features through multi-modal alignment and cross-sensor matching. We provide a comprehensive analysis of multi-sensor transferability, and validate our method on various datasets and in the real-world pouring task. Experimental results show that our method outperforms existing methods, exhibits outstanding static and dynamic perception capabilities across various sensors.

📄 PDF Abstract BibTeX arXiv:2502.12191

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningTransfer Learning

Similar Papers 제목 키워드 기반

AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception

2026-02-10 · Ruoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou 외 arxiv

Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties as well as force dynamics. Although optical tactile senso…

Representation Learning

Dynamic in Static: Hybrid Visual Correspondence for Self-Supervised Video Object Segmentation

2024-04-21 · Gensheng Pei, Yazhou Yao, Jianbo Jiao, Wenguan Wang 외

Conventional video object segmentation (VOS) methods usually necessitate a substantial volume of pixel-level annotated video data for fully supervised learning. In this paper, we present HVC, a \textbf{h}ybrid static-dyn…

Semantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

WildPose: A Unified Framework for Robust Pose Estimation in the Wild

2026-05-12 · Jianhao Zheng, Liyuan Zhu, Zihan Zhu, Iro Armeni arxiv

Estimating camera pose in dynamic environments is a critical challenge, as most visual SLAM and SfM methods assume static scenes. While recent dynamic-aware methods exist, they are often not unified: semantic-based appro…

Pose Estimation

Universal Beta Splatting

2025-09-30 · Rong Liu, Zhongpai Gao, Benjamin Planche, Meida Chen 외 arxiv

We introduce Universal Beta Splatting (UBS), a unified framework that generalizes 3D Gaussian Splatting to N-dimensional anisotropic Beta kernels for explicit radiance field rendering. Unlike fixed Gaussian primitives, B…

EyeWorld: A Generative World Model of Ocular State and Dynamics

2026-03-14 · Ziyu Gao, Xinyuan Wu, Xiaolan Chen, Zhuoran Liu 외 arxiv

Ophthalmic decision-making depends on subtle lesion-scale cues interpreted across multimodal imaging and over time, yet most medical foundation models remain static and degrade under modality and acquisition shifts. Here…

Representation Learning