paper-with-me

홈 › Papers

Cross-Modal Visuo-Tactile Object Perception

2026-04-02 · Anirvan Dutta, Simone Tasciotti, Claudia Cusseddu, Ang Li, Panayiota Poirazi, Julijana Gjorgjieva, Etienne Burdet, Patrick van der Smagt, Mohsen Kaboli arxiv

Estimating physical properties is critical for safe and efficient autonomous robotic manipulation, particularly during contact-rich interactions. In such settings, vision and tactile sensing provide complementary information about object geometry, pose, inertia, stiffness, and contact dynamics, such as stick-slip behavior. However, these properties are only indirectly observable and cannot always be modeled precisely (e.g., deformation in non-rigid objects coupled with nonlinear contact friction), making the estimation problem inherently complex and requiring sustained exploitation of visuo-tactile sensory information during action. Existing visuo-tactile perception frameworks have primarily emphasized forceful sensor fusion or static cross-modal alignment, with limited consideration of how uncertainty and beliefs about object properties evolve over time. Inspired by human multi-sensory perception and active inference, we propose the Cross-Modal Latent Filter (CMLF) to learn a structured, causal latent state-space of physical object properties. CMLF supports bidirectional transfer of cross-modal priors between vision and touch and integrates sensory evidence through a Bayesian inference process that evolves over time. Real-world robotic experiments demonstrate that CMLF improves the efficiency and robustness of latent physical properties estimation under uncertainty compared to baseline approaches. Beyond performance gains, the model exhibits perceptual coupling phenomena analogous to those observed in humans, including susceptibility to cross-modal illusions and similar trajectories in learning cross-sensory associations. Together, these results constitutes a significant step toward generalizable, robust and physically consistent cross-modal integration for robotic multi-sensory perception.

📄 PDF Abstract BibTeX arXiv:2604.02108

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian Inference

Similar Papers 제목 키워드 기반

Universal Visuo-Tactile Video Understanding for Embodied Interaction

2025-05-28 · Yifan Xie, Mingyang Li, Shoujie Li, Xingting Li 외

Tactile perception is essential for embodied agents to understand physical attributes of objects that cannot be determined through visual inspection alone. While existing approaches have made progress in visual and langu…

FrictionLarge Language ModelText GenerationVideo Understanding

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

2025-02-15 · Ruoxuan Feng, Jiangyu Hu, Wenke Xia, Tianci Gao 외

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into rob…

Representation LearningTransfer Learning

Active Cross-Modal Visuo-Tactile Perception of Deformable Linear Objects

2026-01-20 · Raffaele Mazza, Ciro Natale, Pietro Falco arxiv

This paper presents a novel cross-modal visuo-tactile perception framework for the 3D shape reconstruction of deformable linear objects (DLOs), with a specific focus on cables subject to severe visual occlusions. Unlike …

3D Shape ReconstructionInstance SegmentationPoint Clouds

ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

2025-05-26 · Fotios Lygerakis, Ozan Özdenizci, Elmar Rückert

Tactile sensing provides local essential information that is complementary to visual perception, such as texture, compliance, and force. Despite recent advances in visuotactile representation learning, challenges remain …

cross-modal alignmentPositionRepresentation LearningRobotic Grasping+3

A Transfer Learning Approach to Cross-Modal Object Recognition: From Visual Observation to Robotic Haptic Exploration

2020-01-18 · Pietro Falco, Shuang Lu, Ciro Natale, Salvatore Pirozzi 외

In this work, we introduce the problem of cross-modal visuo-tactile object recognition with robotic active exploration. With this term, we mean that the robot observes a set of objects with visual perception and, later o…

ObjectObject RecognitionTransfer Learning