paper-with-me

홈 › Papers

IGen: Scalable Data Generation for Robot Learning from Open-World Images

2025-12-01 · Chenghao Gu, Haolan Kang, Junchao Lin, Jinghe Wang, Duo Wu, Shuzhao Xie, Fanding Huang, Junchen Ge, Ziyang Gong, Letian Li, Hongying Zheng, Changwei Lv, Zhi Wang arxiv

The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-intensive and often limited to specific environments. In contrast, open-world images capture a vast diversity of real-world scenes that naturally align with robotic manipulation tasks, offering a promising avenue for low-cost, large-scale robot data acquisition. Despite this potential, the lack of associated robot actions hinders the practical use of open-world images for robot learning, leaving this rich visual resource largely unexploited. To bridge this gap, we propose IGen, a framework that scalably generates realistic visual observations and executable actions from open-world images. IGen first converts unstructured 2D pixels into structured 3D scene representations suitable for scene understanding and manipulation. It then leverages the reasoning capabilities of vision-language models to transform scene-specific task instructions into high-level plans and generate low-level actions as SE(3) end-effector pose sequences. From these poses, it synthesizes dynamic scene evolution and renders temporally coherent visual observations. Experiments validate the high quality of visuomotor data generated by IGen, and show that policies trained solely on IGen-synthesized data achieve performance comparable to those trained on real-world data. This highlights the potential of IGen to support scalable data generation from open-world images for generalist robotic policy training.

📄 PDF Abstract BibTeX arXiv:2512.01773

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

Rethinking Video Generation Model for the Embodied World

2026-01-21 · Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li 외 arxiv

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synt…

Video Generation

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

2025-03-09 · AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen 외

We explore how scalable robot data can address real-world challenges for generalized robotic manipulation. Introducing AgiBot World, a large-scale platform comprising over 1 million trajectories across 217 tasks in five …

Robotic Agentic Platform for Intelligent Electric Vehicle Disassembly

2026-03-19 · Zachary Allen, Max Conway, Lyle Antieau, Allen Ponraj 외 arxiv

Electric vehicles (EV) create an urgent need for scalable battery recycling, yet disassembly of EV battery packs remains largely manual due to high design variability. We present our Robotic Agentic Platform for Intellig…

Object Detection

Design of scalable orthogonal digital encoding architecture for large-area flexible tactile sensing in robotics

2025-09-13 · Weijie Liu, Ziyi Qiu, Shihang Wang, Deqing Mei 외 arxiv

Human-like embodied tactile perception is crucial for the next-generation intelligent robotics. Achieving large-area, full-body soft coverage with high sensitivity and rapid response, akin to human skin, remains a formid…

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

2026-06-15 · Jie Zhang, Xiaoyue Chen, Anzhe Chen, Dayiheng Liu 외 arxiv

We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual trajectories from curre…

Synthetic Data GenerationAutonomous DrivingVideo Generation