paper-with-me

Papers

GEM: Generative Supervision Helps Embodied Intelligence

2026-05-27 · Ruowen Zhao, Bangguo Li, Zuyan Liu, Yinan Liang, Junliang Ye, Fangfu Liu, Diankun Wu, Zhengyi Wang, Xumin Yu, Yongming Rao, Han Hu, Jun Zhu arxiv

Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a significant gap remains between the high-level semantic focus of standard text-guided pre-training paradigms and the low-level spatial and physical knowledge critical for execution in embodied environments. In this paper, we introduce GEM, a Generative-supervised Embodied vision-language Model designed to bridge this divide. We propose integrating a depth map generation task directly into the VLM pre-training phase. By training this generative objective jointly with the main model, we observe substantial improvements in embodied intelligence, significantly enhancing both semantic understanding and physical operation capabilities. To support this paradigm, we curate and release GEM-4M, a comprehensive large-scale dataset featuring a mixture of grounding, reasoning, and planning data paired with high-quality depth supervision. Extensive experiments demonstrate that GEM achieves state-of-the-art results across diverse embodied benchmarks. Furthermore, our deployed action model, GEM-VLA, exhibits vastly superior task execution abilities in both simulation environments and real-world evaluations. Code, models, and datasets are available at https://zhaorw02.github.io/GEM/

📄 PDF Abstract BibTeX arXiv:2605.28548

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence

2025-06-12 · Xinjie Wang, Liu Liu, Yu Cao, Ruiqi Wu 외

Constructing a physically realistic and accurately scaled simulated 3D world is crucial for the training and evaluation of embodied intelligence tasks. The diversity, realism, low cost accessibility and affordability of …

Image to 3DLayout GenerationScene GenerationText to 3D+1

GridToPix: Training Embodied Agents with Minimal Supervision

2021-04-14 · ICCV 2021 10 · Unnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi 외

While deep reinforcement learning (RL) promises freedom from hand-labeled data, great successes, especially for Embodied AI, require significant work to create supervision via carefully shaped rewards. Indeed, without sh…

Deep Reinforcement LearningPointGoal NavigationReinforcement Learning (RL)Task 2

Advances in Embodied Navigation Using Large Language Models: A Survey

2023-11-01 · Jinzhou Lin, Han Gao, Xuxiang Feng, Rongtao Xu 외

In recent years, the rapid advancement of Large Language Models (LLMs) such as the Generative Pre-trained Transformer (GPT) has attracted increasing attention due to their potential in a variety of practical applications…

Decision Making

Towards a Data Flywheel for Embodied Intelligence in Logistics

2026-06-04 · Anlan Yu, Zaishu Chen, Zhiqing Hong, Daqing Zhang arxiv

Embodied intelligence is moving from laboratory demonstrations toward industrial deployment, with the logistics industry serving as a key application scenario. Learning-based policies offer a promising path beyond tradit…

How Intelligent is your Intelligent Robot?

2017-12-24 · Alan F. T. Winfield

How intelligent is robot A compared with robot B? And how intelligent are robots A and B compared with animals (or plants) X and Y? These are both interesting and deeply challenging questions. In this paper we address th…