paper-with-me

Papers

Self-Corrected Multimodal Large Language Model for End-to-End Robot Manipulation

2024-05-27 · Jiaming Liu, Chenxuan Li, Guanqun Wang, Lily Lee, Kaichen Zhou, Sixiang Chen, Chuyan Xiong, Jiaxin Ge, Renrui Zhang, Shanghang Zhang

Robot manipulation policies have shown unsatisfactory action performance when confronted with novel task or object instances. Hence, the capability to automatically detect and self-correct failure action is essential for a practical robotic system. Recently, Multimodal Large Language Models (MLLMs) have shown promise in visual instruction following and demonstrated strong reasoning abilities in various tasks. To unleash general MLLMs as an end-to-end robotic agent, we introduce a Self-Corrected (SC)-MLLM, equipping our model not only to predict end-effector poses but also to autonomously recognize and correct failure actions. Specifically, we first conduct parameter-efficient fine-tuning to empower MLLM with pose prediction ability, which is reframed as a language modeling problem. When facing execution failures, our model learns to identify low-level action error causes (i.e., position and rotation errors) and adaptively seeks prompt feedback from experts. Based on the feedback, SC-MLLM rethinks the current failure scene and generates the corrected actions. Furthermore, we design a continuous policy learning method for successfully corrected samples, enhancing the model's adaptability to the current scene configuration and reducing the frequency of expert intervention. To evaluate our SC-MLLM, we conduct extensive experiments in both simulation and real-world settings. SC-MLLM agent significantly improve manipulation accuracy compared to previous state-of-the-art robotic MLLM (ManipLLM), increasing from 57\% to 79\% on seen object categories and from 47\% to 69\% on unseen novel categories.

📄 PDF Abstract BibTeX arXiv:2405.17418

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Modelparameter-efficient fine-tuningPose PredictionRobot Manipulationvisual instruction following

Similar Papers 제목 키워드 기반

Sensorimotor features of self-awareness in multimodal large language models

2025-05-25 · Iñaki Dellibarda Varela, Pablo Romero-Sorozabal, Diego Torricelli, Gabriel Delgado-Oleas 외

Self-awareness - the ability to distinguish oneself from the surrounding environment - underpins intelligent, autonomous behavior. Recent advances in AI achieve human-like performance in tasks integrating multimodal info…

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

2025-04-03 · Xiaofeng Han, Shunpeng Chen, Zenghuang Fu, Zhe Feng 외

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We systematically review the applications of multimodal fusion in key robotic vision tasks, includin…

3D Object Detectioncross-modal alignmentDomain Adaptationobject-detection+5

e-SNLI-VE: Corrected Visual-Textual Entailment with Natural Language Explanations

2020-04-07 · Virginie Do, Oana-Maria Camburu, Zeynep Akata, Thomas Lukasiewicz

The recently proposed SNLI-VE corpus for recognising visual-textual entailment is a large, real-world dataset for fine-grained multimodal reasoning. However, the automatic way in which SNLI-VE has been assembled (via com…

Multimodal ReasoningNatural Language Inference

CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation

2023-06-17 · Xiwen Liang, Liang Ma, Shanshan Guo, Jianhua Han 외

Understanding and following natural language instructions while navigating through complex, real-world environments poses a significant challenge for general-purpose robots. These environments often include obstacles and…

Decision MakingInstruction FollowingLanguage ModellingLarge Language Model+2

OctoSense: Self-Supervised Learning for Multimodal Robot Perception

2026-06-25 · Anthony Bisulco, Jeremy Wang, Kostas Daniilidis, Randall Balestriero 외 arxiv

We present OctoSense, an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corrected global positioning system, and proprioception (CAN bus data fro…

Self-Supervised LearningSemantic Segmentation