paper-with-me

홈 › Papers

UMIGen: A Unified Framework for Egocentric Point Cloud Generation and Cross-Embodiment Robotic Imitation Learning

2025-11-12 · Yan Huang, Shoujie Li, Xingting Li, Wenbo Ding arxiv

Data-driven robotic learning faces an obvious dilemma: robust policies demand large-scale, high-quality demonstration data, yet collecting such data remains a major challenge owing to high operational costs, dependence on specialized hardware, and the limited spatial generalization capability of current methods. The Universal Manipulation Interface (UMI) relaxes the strict hardware requirements for data collection, but it is restricted to capturing only RGB images of a scene and omits the 3D geometric information on which many tasks rely. Inspired by DemoGen, we propose UMIGen, a unified framework that consists of two key components: (1) Cloud-UMI, a handheld data collection device that requires no visual SLAM and simultaneously records point cloud observation-action pairs; and (2) a visibility-aware optimization mechanism that extends the DemoGen pipeline to egocentric 3D observations by generating only points within the camera's field of view. These two components enable efficient data generation that aligns with real egocentric observations and can be directly transferred across different robot embodiments without any post-processing. Experiments in both simulated and real-world settings demonstrate that UMIGen supports strong cross-embodiment generalization and accelerates data collection in diverse manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2511.09302

Code (0)

등록된 구현이 없습니다.

Tasks

Point Cloud Generation

Similar Papers 제목 키워드 기반

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

2025-08-05 · Xiaoqi Dong, Xiangyu Zhou, Nicholas Evans, Yujia Lin arxiv

Text-to-Image (T2I) generation has made significant advancements with diffusion models, yet challenges persist in handling complex instructions, ensuring fine-grained content control, and maintaining deep semantic consis…

Text-to-Image GenerationInstruction Following

Egocentric Action-aware Inertial Localization in Point Clouds

2025-05-20 · Mingfang Zhang, Ryo Yonetani, Yifei HUANG, Liangyang Ouyang 외

This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individu…

Action Recognition

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

2026-06-18 · Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le 외 arxiv

Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a …

Representation LearningAction SegmentationAction RecognitionVideo Retrieval

FRAME: Fast and Robust Autonomous 3D point cloud Map-merging for Egocentric multi-robot exploration

2023-01-22 · Nikolaos Stathoulopoulos, Anton Koval, Ali-akbar Agha-mohammadi, George Nikolakopoulos

This article presents a 3D point cloud map-merging framework for egocentric heterogeneous multi-robot exploration, based on overlap detection and alignment, that is independent of a manual initial guess or prior knowledg…

Point Cloud Registration

LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World

2026-05-06 · Nan Yang, Julian Straub, Fan Zhang, Richard Newcombe 외 arxiv

Tracking 3D human motion from egocentric multi-camera headset is challenged by severe egomotion, partial visibility or occlusions and lack of training data. Existing methods designed for monocular video often require sta…