paper-with-me

홈 › Papers

SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception

2025-12-01 · Gurmeher Khurana, Lan Wei, Dandan Zhang arxiv

Contact-rich robotic manipulation requires representations that encode local geometry. Vision provides global context but lacks direct measurements of properties such as texture and hardness, whereas touch supplies these cues. Modern visuo-tactile sensors capture both modalities in a single fused image, yielding intrinsically aligned inputs that are well suited to manipulation tasks requiring visual and tactile information. Most self-supervised learning (SSL) frameworks, however, compress feature maps into a global vector, discarding spatial structure and misaligning with the needs of manipulation. To address this, we propose SARL, a spatially-aware SSL framework that augments the Bootstrap Your Own Latent (BYOL) architecture with three map-level objectives, including Saliency Alignment (SAL), Patch-Prototype Distribution Alignment (PPDA), and Region Affinity Matching (RAM), to keep attentional focus, part composition, and geometric relations consistent across views. These losses act on intermediate feature maps, complementing the global objective. SARL consistently outperforms nine SSL baselines across six downstream tasks with fused visual-tactile data. On the geometry-sensitive edge-pose regression task, SARL achieves a Mean Absolute Error (MAE) of 0.3955, a 30% relative improvement over the next-best SSL method (0.5682 MAE) and approaching the supervised upper bound. These findings indicate that, for fused visual-tactile data, the most effective signal is structured spatial equivariance, in which features vary predictably with object geometry, which enables more capable robotic perception.

📄 PDF Abstract BibTeX arXiv:2512.01908

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningRepresentation Learning

Similar Papers 제목 키워드 기반

SARL*: Deep Reinforcement Learning based Human-Aware Navigation for Mobile Robot in Indoor Environments

2020-01-20 · Keyu Li, Yangxin Xu, Jiankun Wang, Max Q.-H. Meng

In a human-robot coexisting environment, reaching the goal position safely and efficiently is essential for a mobile service robot. In this paper, we present an advanced version of the Socially Attentive Reinforcement Le…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

ViSaRL: Visual Reinforcement Learning Guided by Human Saliency

2024-03-16 · Anthony Liang, Jesse Thomason, Erdem Biyik

Training robots to perform complex control tasks from high-dimensional pixel input using reinforcement learning (RL) is sample-inefficient, because image observations are comprised primarily of task-irrelevant informatio…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Robot Manipulation

SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology

2026-03-30 · Yifan Wang, Bolian Li, David Cho, Ruqi Zhang 외 arxiv

Reinforcement learning is critical to improving large reasoning models, but its success relies heavily on verifiable rewards (RLVR), making it hard to use in open-ended domains where correctness is ambiguous and cannot b…

Reinforcement Learning

Self-Supervised Spatially Variant PSF Estimation for Aberration-Aware Depth-from-Defocus

2024-02-28 · Zhuofeng Wu, Yusuke Monno, Masatoshi Okutomi

In this paper, we address the task of aberration-aware depth-from-defocus (DfD), which takes account of spatially variant point spread functions (PSFs) of a real camera. To effectively obtain the spatially variant PSFs o…

Depth EstimationSelf-Supervised Learning

Learning Representations by Predicting Bags of Visual Words

2020-02-27 · CVPR 2020 6 · Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez 외

Self-supervised representation learning targets to learn convnet-based image representations from unlabeled data. Inspired by the success of NLP methods in this area, in this work we propose a self-supervised approach ba…

Representation Learning