paper-with-me

홈 › Papers

The Treachery of Images: Bayesian Scene Keypoints for Deep Policy Learning in Robotic Manipulation

2023-05-08 · Jan Ole von Hartz, Eugenio Chisari, Tim Welschehold, Wolfram Burgard, Joschka Boedecker, Abhinav Valada

In policy learning for robotic manipulation, sample efficiency is of paramount importance. Thus, learning and extracting more compact representations from camera observations is a promising avenue. However, current methods often assume full observability of the scene and struggle with scale invariance. In many tasks and settings, this assumption does not hold as objects in the scene are often occluded or lie outside the field of view of the camera, rendering the camera observation ambiguous with regard to their location. To tackle this problem, we present BASK, a Bayesian approach to tracking scale-invariant keypoints over time. Our approach successfully resolves inherent ambiguities in images, enabling keypoint tracking on symmetrical objects and occluded and out-of-view objects. We employ our method to learn challenging multi-object robot manipulation tasks from wrist camera observations and demonstrate superior utility for policy learning compared to other representation learning techniques. Furthermore, we show outstanding robustness towards disturbances such as clutter, occlusions, and noisy depth measurements, as well as generalization to unseen objects both in simulation and real-world robotic experiments.

📄 PDF Abstract BibTeX arXiv:2305.04718

Code (1)

robot-learning-freiburg/bask 공식 구현 pytorch

Tasks

Representation LearningRobot Manipulation

Similar Papers 제목 키워드 기반

A Bag of Words Approach for Semantic Segmentation of Monitored Scenes

2013-05-14 · Wassim Bouachir, Atousa Torabi, Guillaume-Alexandre Bilodeau, Pascal Blais

This paper proposes a semantic segmentation method for outdoor scenes captured by a surveillance camera. Our algorithm classifies each perceptually homogenous region as one of the predefined classes learned from a collec…

SegmentationSemantic Segmentation

Self-Supervised Learning of Multi-Object Keypoints for Robotic Manipulation

2022-05-17 · Jan Ole von Hartz, Eugenio Chisari, Tim Welschehold, Abhinav Valada

In recent years, policy learning methods using either reinforcement or imitation have made significant progress. However, both techniques still suffer from being computationally expensive and requiring large amounts of t…

Representation LearningRobot ManipulationSelf-Supervised Learning

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning

2026-06-17 · Jisoo Kim, Sangwon Baik, Taeksoo Kim, Sungjoo Kim 외 arxiv

We present ZeroDex, a zero-shot framework for long-horizon dexterous manipulation that grounds language instructions into executable 3D task plans from calibrated multi-view RGB images. Rather than training an end-to-end…

Correspondence-Oriented Imitation Learning: Flexible Visuomotor Control with 3D Conditioning

2025-12-05 · Yunhao Cao, Zubin Bhaumik, Jessie Jia, Xingyi He 외 arxiv

We introduce Correspondence-Oriented Imitation Learning (COIL), a conditional policy learning framework for visuomotor control with a flexible task representation in 3D. At the core of our approach, each task is defined …

Deep Transformer Network for Monocular Pose Estimation of Ship-Based UAV

2024-06-13 · Maneesha Wickramasuriya, Taeyoung Lee, Murray Snyder

This paper introduces a deep transformer network for estimating the relative 6D pose of a Unmanned Aerial Vehicle (UAV) with respect to a ship using monocular images. A synthetic dataset of ship images is created and ann…

Pose EstimationPosition