paper-with-me

홈 › Papers

GameVLM: A Decision-making Framework for Robotic Task Planning Based on Visual Language Models and Zero-sum Games

2024-05-22 · Aoran Mei, Jianhua Wang, Guo-Niu Zhu, Zhongxue Gan

With their prominent scene understanding and reasoning capabilities, pre-trained visual-language models (VLMs) such as GPT-4V have attracted increasing attention in robotic task planning. Compared with traditional task planning strategies, VLMs are strong in multimodal information parsing and code generation and show remarkable efficiency. Although VLMs demonstrate great potential in robotic task planning, they suffer from challenges like hallucination, semantic complexity, and limited context. To handle such issues, this paper proposes a multi-agent framework, i.e., GameVLM, to enhance the decision-making process in robotic task planning. In this study, VLM-based decision and expert agents are presented to conduct the task planning. Specifically, decision agents are used to plan the task, and the expert agent is employed to evaluate these task plans. Zero-sum game theory is introduced to resolve inconsistencies among different agents and determine the optimal solution. Experimental results on real robots demonstrate the efficacy of the proposed framework, with an average success rate of 83.3%.

📄 PDF Abstract BibTeX arXiv:2405.13751

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationDecision MakingHallucinationScene UnderstandingTask Planning

Similar Papers 제목 키워드 기반

On-Device Robotic Planning: Eliminating Inference Redundancy for Efficient Decision-Making

2026-05-29 · Joonhee Lee, Hyunseung Shin, Hyunmi Kim, Pei Zhang 외 arxiv

Reasoning-based robotic policies using large language and vision-language models achieve strong semantic planning capabilities but mostly suffer from a high inference latency that limits practical real-time deployment. I…

Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems

2026-05-09 · Jiaxin Liu, Peng Yang, Yuping Li, Xinyue Xie arxiv

Driven by the rapid expansion of e-commerce and small-batch production, the size of the intralogistics load unit of finished goods, semi-finished goods and raw materials is steadily shrinking. Totes are gradually replaci…

Multi-agent Reinforcement Learning

Distortion-Resilient Robotic Imitation Learning for Autonomous Cable Routing

2026-06-10 · Hao Wang, Fu-Zhao Ou, Shiqi Wang, Zhaolin Wan 외 arxiv

The rapid development of intelligent control methodologies has endowed robots with powerful autonomous intelligence. Cable routing, a ubiquitous foundational task in industry, provides a rigorous benchmark for robotic de…

Image Quality Assessment

Reward-Centered ReST-MCTS: A Robust Decision-Making Framework for Robotic Manipulation in High Uncertainty Environments

2025-03-07 · Xibai Wang

Monte Carlo Tree Search (MCTS) has emerged as a powerful tool for decision-making in robotics, enabling efficient exploration of large search spaces. However, traditional MCTS methods struggle in environments characteriz…

Decision MakingEfficient Exploration

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

2025-06-14 · Wenbo Li, Shiyi Wang, Yiteng Chen, Huiping Zhuang 외

Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediat…

Decision MakingQuestion AnsweringVisual Question Answering