paper-with-me

Papers

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

2026-07-02 · Yuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang, Zhanyu Ma arxiv

Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance. We present MMBench-Live, a continuously evolving multimodal benchmark built by a multi-agent-driven automated pipeline. Our framework treats benchmark evolution as task-guided dataset construction, integrating structured benchmark specification, feedback-controlled real-time data acquisition, and verifiable QA generation with executable reasoning. To maintain cross-version comparability, we introduce a distribution-consistent update strategy that extracts task-related visual patterns from the original benchmark to guide data collection and filtering. Instantiated from MMBench, MMBench-Live contains 5.9K newly generated evaluation instances with a high answer correctness rate, while each update costs about USD 30 and takes 1-2 hours. Extensive evaluations show that MMBench-Live preserves stable model rankings, maintains semantic alignment with the original benchmark, and exhibits weaker contamination-related memorization signals, suggesting a practical and scalable paradigm for sustainable multimodal benchmark evolution. The project is available at https://github.com/PRIS-CV/MMBench-Live.

📄 PDF Abstract BibTeX arXiv:2607.01813

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FewMMBench: A Benchmark for Multimodal Few-Shot Learning

2026-02-25 · Mustafa Dogan, Ilker Kesen, Iacer Calixto, Aykut Erdem 외 arxiv

As multimodal large language models (MLLMs) advance in handling interleaved image-text data, assessing their few-shot learning capabilities remains an open challenge. In this paper, we introduce FewMMBench, a comprehensi…

Few-Shot Learning

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

2026-05-15 · Huacan Chai, Yukai Wang, Yingxuan Yang, Dan Peng 외 arxiv

Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use evidence distributed across independently originated sources. We argue…

Multimodal Reasoning

AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy

2025-09-29 · Jinghang Shi, Xiaoyu Tang, Yang Huang, Yuyang Li 외 arxiv

Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus on general multimodal capabilities but f…

INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance

2024-06-13 · Chenwei Lin, Hanjia Lyu, Xian Xu, Jiebo Luo

Large Vision-Language Models (LVLMs) have demonstrated outstanding performance in various general multimodal applications such as image recognition and visual reasoning, and have also shown promising potential in special…

Multiple-choiceVisual Reasoning

Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM

2025-03-18 · Xinyu Fang, Zhijian Chen, Kai Lan, Shengyuan Ding 외

Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) have been extensively evaluated for their …