paper-with-me

홈 › Papers

Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface

2025-12-22 · Yujie Zhao, Hongwei Fan, Di Chen, Shengcong Chen, Liliang Chen, Xiaoqi Li, Guanghui Ren, Hao Dong arxiv

Recent progress in robot learning has been driven by large-scale datasets and powerful visuomotor policy architectures, yet policy robustness remains limited by the substantial cost of collecting diverse demonstrations, particularly for spatial generalization in manipulation tasks. To reduce repetitive data collection, we present Real2Edit2Real, a framework that generates new demonstrations by bridging 3D editability with 2D visual data through a 3D control interface. Our approach first reconstructs scene geometry from multi-view RGB observations with a metric-scale 3D reconstruction model. Based on the reconstructed geometry, we perform depth-reliable 3D editing on point clouds to generate new manipulation trajectories while geometrically correcting the robot poses to recover physically consistent depth, which serves as a reliable condition for synthesizing new demonstrations. Finally, we propose a multi-conditional video generation model guided by depth as the primary control signal, together with action, edge, and ray maps, to synthesize spatially augmented multi-view manipulation videos. Experiments on four real-world manipulation tasks demonstrate that policies trained on data generated from only 1-5 source demonstrations can match or outperform those trained on 50 real-world demonstrations, improving data efficiency by up to 10-50x. Moreover, experimental results on height and texture editing demonstrate the framework's flexibility and extensibility, indicating its potential to serve as a unified data generation framework. Project website is https://real2edit2real.github.io/.

📄 PDF Abstract BibTeX arXiv:2512.19402

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionVideo GenerationPoint Clouds

Similar Papers 제목 키워드 기반

RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping

2025-06-10 · Yang Bai, Liudi Yang, George Eskandar, Fengyi Shen 외

Recent advancements in generative models have revolutionized video synthesis and editing. However, the scarcity of diverse, high-quality datasets continues to hinder video-conditioned robotic learning, limiting cross-pla…

Video Editing

RoboTransfer: Geometry-Consistent Video Diffusion for Robotic Visual Policy Transfer

2025-05-29 · Liu Liu, XiaoFeng Wang, Guosheng Zhao, Keyu Li 외

Imitation Learning has become a fundamental approach in robotic manipulation. However, collecting large-scale real-world robot demonstrations is prohibitively expensive. Simulators offer a cost-effective alternative, but…

Imitation LearningVideo Generation

Splat-MOVER: Multi-Stage, Open-Vocabulary Robotic Manipulation via Editable Gaussian Splatting

2024-05-07 · Ola Shorinwa, Johnathan Tucker, Aliyah Smith, Aiden Swann 외

We present Splat-MOVER, a modular robotics stack for open-vocabulary robotic manipulation, which leverages the editability of Gaussian Splatting (GSplat) scene representations to enable multi-stage manipulation tasks. Sp…

Grasp GenerationSimulated Gaussian Manipulation

Task Editing for Generalizable 3D Visuomotor Policy Learning

2026-06-05 · Jian-Jian Jiang, YiHan Yang, Lan Wei, Yuming Luo 외 arxiv

3D visuomotor policies offer a promising direction for complex robotic manipulation, as depth maps and point clouds provide rich geometric information for spatial reasoning. However, their success often depends on large-…

Spatial ReasoningPoint Clouds

ArticuBot: Learning Universal Articulated Object Manipulation Policy via Large Scale Simulation

2025-03-04 · YuFei Wang, Ziyu Wang, Mino Nakura, Pratik Bhowal 외

This paper presents ArticuBot, in which a single learned policy enables a robotics system to open diverse categories of unseen articulated objects in the real world. This task has long been challenging for robotics due t…

Imitation LearningMotion Planning