paper-with-me

홈 › Papers

UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

2026-04-15 · Ziming Wang arxiv

We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation. While UMI enables portable, wrist-mounted data acquisition, its reliance on monocular visual SLAM makes it vulnerable to occlusions, dynamic scenes, and tracking failures, limiting its applicability in real-world environments. UMI-3D addresses these limitations by introducing a lightweight and low-cost LiDAR sensor tightly integrated into the wrist-mounted interface, enabling LiDAR-centric SLAM with accurate metric-scale pose estimation under challenging conditions. We further develop a hardware-synchronized multimodal sensing pipeline and a unified spatiotemporal calibration framework that aligns visual observations with LiDAR point clouds, producing consistent 3D representations of demonstrations. Despite maintaining the original 2D visuomotor policy formulation, UMI-3D significantly improves the quality and reliability of collected data, which directly translates into enhanced policy performance. Extensive real-world experiments demonstrate that UMI-3D not only achieves high success rates on standard manipulation tasks, but also enables learning of tasks that are challenging or infeasible for the original vision-only UMI setup, including large deformable object manipulation and articulated object operation. The system supports an end-to-end pipeline for data acquisition, alignment, training, and deployment, while preserving the portability and accessibility of the original UMI. All hardware and software components are open-sourced to facilitate large-scale data collection and accelerate research in embodied intelligence: \href{https://umi-3d.github.io}{https://umi-3d.github.io}.

📄 PDF Abstract BibTeX arXiv:2604.14089

Code (0)

등록된 구현이 없습니다.

Tasks

Pose EstimationPoint Clouds

Similar Papers 제목 키워드 기반

YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale

2026-06-08 · Takehiko Ohkawa, Jumpei Arima, Yuki Noguchi, Masatoshi Tateno 외 arxiv

We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collection for bimanual dexterous manipulation. While handheld data collecti…

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning

2026-06-04 · Chaoyi Xu, Yixuan Jiang, Jiahui Huan, Yuhui Fu 외 arxiv

Learning dexterous manipulation requires demonstrations that preserve fine hand-object interactions while remaining executable at deployment. Existing pipelines either lose deployable dexterity through retargeting or emb…

DexHiL: A Human-in-the-Loop Framework for Vision-Language-Action Model Post-Training in Dexterous Manipulation

2026-03-10 · Yifan Han, Zhongxi Chen, Yuxuan Zhao, Congsheng Xu 외 arxiv

While Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities in robotic manipulation, deploying them on specific and complex downstream tasks still demands effective post-training. In…

PointAction: 3D Points as Universal Action Representations for Robot Control

2026-06-02 · Mutian Tong, Han Jiang, Qiao Feng, Lingjie Liu 외 arxiv

Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. However, RGB-only video rollouts are not di…

Robot ManipulationVideo GenerationVideo Prediction

Guava: An Effective and Universal Harness for Embodied Manipulation

2026-06-16 · Haowen Liu, Xirui Li, Shaoxiong Yao, Peng Shi 외 arxiv

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-languag…