paper-with-me

홈 › Papers

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

2025-10-20 · Yulin Luo, Chun-Kai Fan, Menghang Dong, Jiayu Shi, Xiangju Mi, Mengdi Zhao, Bo-Wen Zhang, Cheng Chi, Jiaming Liu, Gaole Dai, Rongyu Zhang, Ruichuan An, Kun Wu, Zhengping Che, Shaoxuan Xie, Guocai Yao, Zhongxia Zhao, Pengwei Wang, Guang Liu, Zhongyuan Wang, Tiejun Huang, Shanghang Zhang arxiv

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, where System 2 performs high-level reasoning and System 1 handles low-level control. We refer to System 2 as the embodied brain, the cognitive core for decision-making in manipulation. Although evaluating this embodied brain is crucial, existing benchmarks mainly measure execution success or cover only limited aspects of high-level cognition and task realism. We introduce RoboBench, a benchmark for evaluating multimodal large language models (MLLMs) as embodied brains. RoboBench covers five dimensions: Instruction Comprehension, Perception Reasoning, Generalized Planning, Affordance Prediction, and Failure Analysis. It spans 14 capabilities, 25 tasks, and 6,092 QA pairs. To improve realism, it draws from large-scale real robotic data and in-house collection across diverse embodiments, attribute-rich objects, multi-view scenes, and memory-driven navigation. For planning, RoboBench introduces an MLLM-as-world-simulator framework that assesses whether predicted plans can achieve critical object-state changes under physical and visual constraints, enabling more faithful evaluation of long-horizon reasoning than symbolic matching. Experiments on 18 state-of-the-art MLLMs reveal persistent limitations in implicit instruction understanding, spatiotemporal reasoning, cross-scenario planning, fine-grained affordance understanding, and failure diagnosis. We further analyze how embodied cognitive abilities relate to downstream robotic control. RoboBench offers a comprehensive scaffold for quantifying high-level cognition and guiding next-generation MLLMs toward more robust robotic intelligence.

📄 PDF Abstract BibTeX arXiv:2510.17801

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoboBenchMart: Benchmarking Robots in Retail Environment

2025-11-13 · Konstantin Soshin, Alexander Krapukhin, Andrei Spiridonov, Gregorii Bukhtuev 외 arxiv

Most existing robotic manipulation benchmarks focus on tabletop or household scenarios. While these setups have driven impressive progress, it remains unclear whether generalist VLAs that excel there can truly generalize…

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

2025-10-22 · Zhiyuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du 외 arxiv

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet …

Spatial Reasoning

MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?

2024-06-28 · Jinming Li, Yichen Zhu, Zhiyuan Xu, Jindong Gu 외

It is fundamentally challenging for robots to serve as useful assistants in human environments because this requires addressing a spectrum of sub-problems across robotics, including perception, language understanding, re…

Task PlanningVisual Reasoning

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

2023-05-13 · Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang 외

Large models have recently played a dominant role in natural language processing and multimodal vision-language learning. However, their effectiveness in text-related visual tasks remains relatively unexplored. In this p…

Key Information ExtractionNutritionOptical Character RecognitionOptical Character Recognition (OCR)+3

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

2023-06-23 · Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin 외

Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for t…

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model+4