paper-with-me

Papers

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

2025-11-25 · GigaWorld Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jiagang Zhu, Kerui Li, Mengyuan Xu, Qiuping Deng, Siting Wang, Wenkang Qin, Xinze Chen, Xiaofeng Wang, Yankai Wang, Yu Cao, Yifan Chang, Yuan Xu, Yun Ye, Yang Wang, Yukun Zhou, Zhengyuan Zhang, Zhehao Dong, Zheng Zhu arxiv

World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA) learning. GigaWorld-0 integrates two synergistic components: GigaWorld-0-Video, which leverages large-scale video generation to produce diverse, texture-rich, and temporally coherent embodied sequences under fine-grained control of appearance, camera viewpoint, and action semantics; and GigaWorld-0-3D, which combines 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning to ensure geometric consistency and physical realism. Their joint optimization enables the scalable synthesis of embodied interaction data that is visually compelling, spatially coherent, physically plausible, and instruction-aligned. Training at scale is made feasible through our efficient GigaTrain framework, which exploits FP8-precision and sparse attention to drastically reduce memory and compute requirements. We conduct comprehensive evaluations showing that GigaWorld-0 generates high-quality, diverse, and controllable data across multiple dimensions. Critically, VLA model (e.g., GigaBrain-0) trained on GigaWorld-0-generated data achieve strong real-world performance, significantly improving generalization and task success on physical robots without any real-world interaction during training.

📄 PDF Abstract BibTeX arXiv:2511.19861

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationMotion Planning

Similar Papers 제목 키워드 기반

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

2026-07-15 · GigaWorld Team, Angen Ye, Angyuan Ma, Boyuan Wang 외 arxiv

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a…

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

2026-07-02 · GigaWorld Team, Angyuan Ma, Boyuan Wang, Bohan Li 외 hf

Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by ha…

GigaWorld-Policy: An Efficient Action-Centered World--Action Model

2026-03-18 · Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang 외 arxiv

World-Action Models (WAM) initialized from pre-trained video generation backbones have demonstrated remarkable potential for robot policy learning. However, existing approaches face two critical bottlenecks that hinder p…

Video Generation

Embodied AI: From LLMs to World Models

2025-09-24 · Tongtong Feng, Xin Wang, Yu-Gang Jiang, Wenwu Zhu arxiv

Embodied Artificial Intelligence (AI) is an intelligent system paradigm for achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications and driving the evolution from cyberspace t…

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling

2026-05-03 · Peiyan Tu, Hanxin Zhu, Jingwen Sun, Shaojie Ren 외 arxiv

Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However, existing robot data are typically capt…

Spatial ReasoningDecision Making