paper-with-me

홈 › Papers

GigaBrain-0: A World Model-Powered Vision-Language-Action Model

2025-10-22 · GigaBrain Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jie Li, Jiagang Zhu, Lv Feng, Peng Li, Qiuping Deng, Runqi Ouyang, Wenkang Qin, Xinze Chen, Xiaofeng Wang, Yang Wang, Yifan Li, Yilong Li, Yiran Ding, Yuan Xu, Yun Ye, Yukun Zhou, Zhehao Dong, Zhenan Wang, Zhichao Liu, Zheng Zhu arxiv

Training Vision-Language-Action (VLA) models for generalist robots typically requires large-scale real-world robot data, which is expensive and time-consuming to collect. The inefficiency of physical data collection severely limits the scalability, and generalization capacity of current VLA systems. To address this challenge, we introduce GigaBrain-0, a novel VLA foundation model empowered by world model-generated data (e.g., video generation, real2real transfer, human transfer, view transfer, sim2real transfer data). By leveraging world models to generate diverse data at scale, GigaBrain-0 significantly reduces reliance on real robot data while improving cross-task generalization. Our approach further improves policy robustness through RGBD input modeling and embodied Chain-of-Thought (CoT) supervision, enabling the model to reason about spatial geometry, object states, and long-horizon dependencies during task execution. This leads to substantial gains in real-world performance on dexterous, long-horizon, and mobile manipulation tasks. Extensive experiments demonstrate that GigaBrain-0 achieves superior generalization across variations in appearances (e.g., textures, colors), object placements, and camera viewpoints. Additionally, we present GigaBrain-0-Small, an optimized lightweight variant designed to run efficiently on devices such as the NVIDIA Jetson AGX Orin.

📄 PDF Abstract BibTeX arXiv:2510.19430

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning

2026-02-12 · GigaBrain Team, Boyuan Wang, Bohan Li, Chaojun Ni 외 arxiv

Vision-language-action (VLA) models that directly predict multi-step action chunks from current observations face inherent limitations due to constrained scene understanding and weak future anticipation capabilities. In …

Reinforcement LearningScene Understanding

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

2026-08-16 · GigaBrain Team, Angen Ye, Axiang Sun, Can Jin 외 hf

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question wh…

Instruction Following

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

2025-11-25 · GigaWorld Team, Angen Ye, Boyuan Wang, Chaojun Ni 외 arxiv

World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Languag…

Video GenerationMotion Planning

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

2026-07-02 · GigaWorld Team, Angyuan Ma, Boyuan Wang, Bohan Li 외 hf

Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by ha…

WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation

2025-03-04 · Dujun Nie, Xianda Guo, Yiqun Duan, Ruijun Zhang 외

Object Goal Navigation-requiring an agent to locate a specific object in an unseen environment-remains a core challenge in embodied AI. Although recent progress in Vision-Language Model (VLM)-based agents has demonstrate…

Hallucination