paper-with-me

홈 › Papers

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

2025-10-15 · Kai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian, Dian Zheng, Hongbo Liu, Jingwen He, Bin Liu, Yu Qiao, Ziwei Liu arxiv

Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. Existing evaluations either treat the two abilities in isolation or overlook tasks that inherently couple them. To address this gap, we present Uni-MMMU, a comprehensive and discipline-aware benchmark that systematically unfolds the bidirectional synergy between generation and understanding across eight reasoning-centric domains, including science, coding, mathematics, and puzzles. Each task is bidirectionally coupled, demanding models to (i) leverage conceptual understanding to guide precise visual synthesis, or (ii) utilize generation as a cognitive scaffold for analytical reasoning. Uni-MMMU incorporates verifiable intermediate reasoning steps, unique ground truths, and a reproducible scoring protocol for both textual and visual outputs. Through extensive evaluation of state-of-the-art unified, generation-only, and understanding-only models, we reveal substantial performance disparities and cross-modal dependencies, offering new insights into when and how these abilities reinforce one another, and establishing a reliable foundation for advancing unified models.

📄 PDF Abstract BibTeX arXiv:2510.13759

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

2023-11-27 · CVPR 2024 1 · Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng 외

We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected m…

Complex Query AnsweringLogical ReasoningVisual Reasoning

CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark

2024-01-22 · Ge Zhang, Xinrun Du, Bei Chen, Yiming Liang 외

As the capabilities of large multimodal models (LMMs) continue to advance, evaluating the performance of LMMs emerges as an increasing need. Additionally, there is an even larger gap in evaluating the advanced knowledge …

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

2024-09-04 · Xiang Yue, Tianyu Zheng, Yuansheng Ni, YuBo Wang 외

This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning c…

Optical Character Recognition (OCR)

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

2026-03-18 · Nahyun Lee, Guijin Son, Hyunwoo Ko, Chanyoung Kim 외 arxiv

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams natively written in Korean, covering nine dis…

JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

2024-10-22 · Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira 외

Accelerating research on Large Multimodal Models (LMMs) in non-English languages is crucial for enhancing user experiences across broader populations. In this paper, we introduce JMMMU (Japanese MMMU), the first large-sc…

Math