paper-with-me

홈 › Papers

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

2024-10-08 · Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, Hanbo Zhang, Minzhao Zhu

We present GR-2, a state-of-the-art generalist robot agent for versatile and generalizable robot manipulation. GR-2 is first pre-trained on a vast number of Internet videos to capture the dynamics of the world. This large-scale pre-training, involving 38 million video clips and over 50 billion tokens, equips GR-2 with the ability to generalize across a wide range of robotic tasks and environments during subsequent policy learning. Following this, GR-2 is fine-tuned for both video generation and action prediction using robot trajectories. It exhibits impressive multi-task learning capabilities, achieving an average success rate of 97.7% across more than 100 tasks. Moreover, GR-2 demonstrates exceptional generalization to new, previously unseen scenarios, including novel backgrounds, environments, objects, and tasks. Notably, GR-2 scales effectively with model size, underscoring its potential for continued growth and application. Project page: \url{https://gr2-manipulation.github.io}.

📄 PDF Abstract BibTeX arXiv:2410.06158

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningRobot ManipulationVideo Generation

Similar Papers 제목 키워드 기반

Text2Action: Generative Adversarial Synthesis from Language to Action

2017-10-15 · Hyemin Ahn, Timothy Ha, Yunho Choi, Hwiyeon Yoo 외

In this paper, we propose a generative model which learns the relationship between language and human action in order to generate a human action sequence given a sentence describing human behavior. The proposed generativ…

DecoderGenerative Adversarial NetworkSentence

BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation

2025-05-19 · Haiquan Wen, Yiwei He, Zhenglin Huang, Tianxiao Li 외

Advances in AI generative models facilitate super-realistic video synthesis, amplifying misinformation risks via social media and eroding trust in digital content. Several research works have explored new deepfake detect…

Binary ClassificationDeepFake DetectionFace SwappingLarge Language Model+3

REST: REtrieve & Self-Train for generative action recognition

2022-09-29 · Adrian Bulat, Enrique Sanchez, Brais Martinez, Georgios Tzimiropoulos

This work is on training a generative action/video recognition model whose output is a free-form action-specific caption describing the video (rather than an action class label). A generative approach has practical advan…

Action RecognitionCaption GenerationContrastive LearningRetrieval+3

VideoNeuMat: Neural Material Extraction from Generative Video Models

2026-02-06 · Bowen Xue, Saeed Hadadan, Zheng Zeng, Fabrice Rousselle 외 arxiv

Creating photorealistic materials for 3D rendering requires exceptional artistic skill. Generative models for materials could help, but are currently limited by the lack of high-quality training data. While recent video …

Large Video Planner Enables Generalizable Robot Control

2025-12-17 · Boyuan Chen, Tianyuan Zhang, Haoran Geng, Caiyi Zhang 외 arxiv

General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundation models by extending multimodal large language models (MLLMs) with action ou…

Instruction FollowingTemporal Sequences