paper-with-me

Papers

Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models

2024-02-22 · Jinyi Liu, Yifu Yuan, Jianye Hao, Fei Ni, Lingzhi Fu, Yibin Chen, Yan Zheng

Recently, there has been considerable attention towards leveraging large language models (LLMs) to enhance decision-making processes. However, aligning the natural language text instructions generated by LLMs with the vectorized operations required for execution presents a significant challenge, often necessitating task-specific details. To circumvent the need for such task-specific granularity, inspired by preference-based policy learning approaches, we investigate the utilization of multimodal LLMs to provide automated preference feedback solely from image inputs to guide decision-making. In this study, we train a multimodal LLM, termed CriticGPT, capable of understanding trajectory videos in robot manipulation tasks, serving as a critic to offer analysis and preference feedback. Subsequently, we validate the effectiveness of preference labels generated by CriticGPT from a reward modeling perspective. Experimental evaluation of the algorithm's preference accuracy demonstrates its effective generalization ability to new tasks. Furthermore, performance on Meta-World tasks reveals that CriticGPT's reward model efficiently guides policy learning, surpassing rewards based on state-of-the-art pre-trained representation models.

📄 PDF Abstract BibTeX arXiv:2402.14245

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingRobot Manipulation

Similar Papers 제목 키워드 기반

Self-Corrected Multimodal Large Language Model for End-to-End Robot Manipulation

2024-05-27 · Jiaming Liu, Chenxuan Li, Guanqun Wang, Lily Lee 외

Robot manipulation policies have shown unsatisfactory action performance when confronted with novel task or object instances. Hence, the capability to automatically detect and self-correct failure action is essential for…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+5

ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

2025-05-28 · Jiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren 외

Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grai…

Contact-rich ManipulationMixture-of-ExpertsVision-Language-Action

Learning Robotic Manipulation Skills Using an Adaptive Force-Impedance Action Space

2021-10-19 · Maximilian Ulmer, Elie Aljalbout, Sascha Schwarz, Sami Haddadin

Intelligent agents must be able to think fast and slow to perform elaborate manipulation tasks. Reinforcement Learning (RL) has led to many promising results on a range of challenging decision-making tasks. However, in r…

Contact-rich ManipulationDecision Makingreinforcement-learningReinforcement Learning+1

Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation

2025-11-13 · Xiangyi Wei, Haotian Zhang, Xinyi Cao, Siyu Xie 외 arxiv

The Vision-Language-Action models (VLA) have achieved significant advances in robotic manipulation recently. However, vision-only VLA models create fundamental limitations, particularly in perceiving interactive and mani…

Audio Generation

Data-Agnostic Robotic Long-Horizon Manipulation with Vision-Language-Guided Closed-Loop Feedback

2025-03-27 · Yuan Meng, Xiangtong Yao, Haihui Ye, Yirui Zhou 외

Recent advances in language-conditioned robotic manipulation have leveraged imitation and reinforcement learning to enable robots to execute tasks from human commands. However, these methods often suffer from limited gen…

Task Planning