paper-with-me

홈 › Papers

Galaxea Open-World Dataset and G0 Dual-System VLA Model

2025-08-30 · Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, Hang Zhao arxiv

We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gathered using a consistent robotic embodiment, paired with precise subtask-level language annotations to facilitate both training and evaluation. Building on this dataset, we introduce G0, a dual-system framework that couples a Vision-Language Model (VLM) for multimodal planning with a Vision-Language-Action (VLA) model for fine-grained execution. G0 is trained using a three-stage curriculum: cross-embodiment pre-training, single-embodiment pre-training, and task-specific post-training. A comprehensive benchmark spanning tabletop manipulation, few-shot learning, and long-horizon mobile manipulation, demonstrates the effectiveness of our approach. In particular, we find that the single-embodiment pre-training stage, together with the Galaxea Open-World Dataset, plays a critical role in achieving strong performance.

📄 PDF Abstract BibTeX arXiv:2509.00576

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Learning

Similar Papers 제목 키워드 기반

AffordTrajDP: Dynamic Affordance-Guided Visuomotor Policy Learning for Robotic Manipulation

2026-08-03 · Gaoyuan Wu, Ziyu Shan, Haoyang Du, Yuyao Jiang 외 arxiv

Affordance-guided imitation learning has shown impressive performance in robotic manipulation tasks by compressing visual perception into task-specific geometric constraints (e.g., fixed contact points). However, the com…

CORAL: Scalable Multi-Task Robot Learning via LoRA Experts

2026-03-10 · Yuankai Luo, Woping Chen, Tong Liang, Zhenguo Li arxiv

Deploying Vision-Language-Action (VLA) models in real-world robotics exposes a core multi-task learning challenge: reconciling task interference in multi-task robotic learning. When multiple tasks are jointly fine-tuned …

Multi-Task Learning

AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models

2026-03-09 · Xiaoquan Sun, Zetian Xu, Chen Cao, Zonghe Liu 외 arxiv

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be improved by robust instruction grounding, a …

Deep Metric Learning for Open World Semantic Segmentation

2021-08-10 · ICCV 2021 10 · Jun Cen, Peng Yun, Junhao Cai, Michael Yu Wang 외

Classical close-set semantic segmentation networks have limited ability to detect out-of-distribution (OOD) objects, which is important for safety-critical applications such as autonomous driving. Incrementally learning …

Autonomous DrivingFew-Shot LearningMetric LearningSegmentation+1

Learngene: From Open-World to Your Learning Task

2021-06-12 · Qiufeng Wang, Xin Geng, Shuxia Lin, Shiyu Xia 외

Although deep learning has made significant progress on fixed large-scale datasets, it typically encounters challenges regarding improperly detecting unknown/unseen classes in the open-world scenario, over-parametrized, …