paper-with-me

홈 › Papers

Vision-Language Foundation Models as Effective Robot Imitators

2023-11-02 · Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Jie Xu, Hongtao Wu, Chilam Cheang, Ya Jing, Weinan Zhang, Huaping Liu, Hang Li, Tao Kong

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on robotics data. To this end, we derive a simple and novel vision-language manipulation framework, dubbed RoboFlamingo, built upon the open-source VLMs, OpenFlamingo. Unlike prior works, RoboFlamingo utilizes pre-trained VLMs for single-step vision-language comprehension, models sequential history information with an explicit policy head, and is slightly fine-tuned by imitation learning only on language-conditioned manipulation datasets. Such a decomposition provides RoboFlamingo the flexibility for open-loop control and deployment on low-performance platforms. By exceeding the state-of-the-art performance with a large margin on the tested benchmark, we show RoboFlamingo can be an effective and competitive alternative to adapt VLMs to robot control. Our extensive experimental results also reveal several interesting conclusions regarding the behavior of different pre-trained VLMs on manipulation tasks. We believe RoboFlamingo has the potential to be a cost-effective and easy-to-use solution for robotics manipulation, empowering everyone with the ability to fine-tune their own robotics policy.

📄 PDF Abstract BibTeX arXiv:2311.01378

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningRobot Manipulation

Similar Papers 제목 키워드 기반

Adaptive t-Momentum-based Optimization for Unknown Ratio of Outliers in Amateur Data in Imitation Learning

2021-08-02 · Wendyam Eric Lionel Ilboudo, Taisuke Kobayashi, Kenji Sugimoto

Behavioral cloning (BC) bears a high potential for safe and direct transfer of human skills to robots. However, demonstrations performed by human operators often contain noise or imperfect behaviors that can affect the e…

Imitation Learning

Evolution of the Science Fiction Writer's Capacity to Imagine the Future

2019-03-13

Drawing upon a body of research on the evolution of creativity, this paper proposes a theory of how, when, and why the forward-thinking story-telling abilities of humans evolved, culminating in the visionary abilities of…

AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents

2024-01-23 · Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas 외

Foundation models that incorporate language, vision, and more recently actions have revolutionized the ability to harness internet scale data to reason about useful tasks. However, one of the key challenges of training e…

Instruction FollowingScene Understanding

Differentiable Robot Rendering

2024-10-17 · Ruoshi Liu, Alper Canberk, Shuran Song, Carl Vondrick

Vision foundation models trained on massive amounts of visual data have shown unprecedented reasoning and planning skills in open-world settings. A key challenge in applying them to robotic tasks is the modality gap betw…

Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review

2025-05-26 · Matthew Lisondra, Beno Benhabib, Goldie Nejat

Rapid advancements in foundation models, including Large Language Models, Vision-Language Models, Multimodal Large Language Models, and Vision-Language-Action Models have opened new avenues for embodied AI in mobile serv…

Decision Making Under UncertaintySensor FusionVision-Language-Action