paper-with-me

홈 › Papers

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents

2025-07-25 · Xuehui Wang, Zhenyu Wu, JingJing Xie, Zichen Ding, Bowen Yang, Zehao Li, Zhaoyang Liu, Qingyun Li, Xuan Dong, Zhe Chen, Weiyun Wang, Xiangyu Zhao, Jixuan Chen, Haodong Duan, Tianbao Xie, Chenyu Yang, Shiqian Su, Yue Yu, Yuan Huang, Yiqian Liu, Xiao Zhang, Yanting Zhang, Xiangyu Yue, Weijie Su, Xizhou Zhu, Wei Shen, Jifeng Dai, Wenhai Wang arxiv

We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of modular frameworks that integrate specialized grounding modules. Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role. More important, task efficiency remains a critically underexplored dimension, and all models suffer from substantial inefficiencies, with excessive redundant steps even when tasks are ultimately completed. The integration of precise localization, effective planning, and early stopping strategies is indispensable to enable truly efficient and scalable GUI automation. Our benchmark code, evaluation data, and running environment will be publicly available at https://github.com/open-compass/MMBench-GUI.

📄 PDF Abstract BibTeX arXiv:2507.19478

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

2026-07-02 · Yuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang 외 arxiv

Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance. We pre…

MMBench: Is Your Multi-modal Model an All-around Player?

2023-07-12 · YuAn Liu, Haodong Duan, Yuanhan Zhang, Bo Li 외

Large vision-language models (VLMs) have recently achieved remarkable progress, exhibiting impressive multimodal perception and reasoning abilities. However, effectively evaluating these large VLMs remains a major challe…

AllInstruction FollowingMultiple-choiceVisual Question Answering

AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models

2024-06-13 · Yuhang Wu, Wenmeng Yu, Yean Cheng, Yan Wang 외

Evaluating the alignment capabilities of large Vision-Language Models (VLMs) is essential for determining their effectiveness as helpful assistants. However, existing benchmarks primarily focus on basic abilities using n…

Multiple-choice

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

2024-06-20 · Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao 외

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative…

FormVideo Understanding

Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM

2025-03-18 · Xinyu Fang, Zhijian Chen, Kai Lan, Shengyuan Ding 외

Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) have been extensively evaluated for their …