paper-with-me

Papers

JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction

2025-12-16 · Atsuyuki Miyai, Shota Onohara, Jeonghun Baek, Kiyoharu Aizawa arxiv

This paper introduces JMMMU-Pro, an image-based Japanese Multi-discipline Multimodal Understanding Benchmark, and Vibe Benchmark Construction, a scalable construction method. Following the evolution from MMMU to MMMU-Pro, JMMMU-Pro extends JMMMU by composing the question image and question text into a single image, thereby creating a benchmark that requires integrated visual-textual understanding through visual perception. To build JMMMU-Pro, we propose Vibe Benchmark Construction, a methodology in which an image generative model (e.g., Nano Banana Pro) produces candidate visual questions, and humans verify the outputs and, when necessary, regenerate with adjusted prompts to ensure quality. By leveraging Nano Banana Pro's highly realistic image generation capabilities and its ability to embed clean Japanese text, we construct a high-quality benchmark at low cost, covering a wide range of background and layout designs. Experimental results show that all open-source LMMs struggle substantially with JMMMU-Pro, underscoring JMMMU-Pro as an important benchmark for guiding future efforts in the open-source community. We believe that JMMMU-Pro provides a more rigorous evaluation tool for assessing the Japanese capabilities of LMMs and that our Vibe Benchmark Construction also offers an efficient guideline for future development of image-based VQA benchmarks.

📄 PDF Abstract BibTeX arXiv:2512.14620

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

2024-10-22 · Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira 외

Accelerating research on Large Multimodal Models (LMMs) in non-English languages is crucial for enhancing user experiences across broader populations. In this paper, we introduce JMMMU (Japanese MMMU), the first large-sc…

Math

Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model

2024-10-30 · Keito Sasagawa, Koki Maeda, Issa Sugiura, Shuhei Kurita 외

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abun…

Language ModelingLanguage Modelling

Evaluating Multimodal Large Language Models on Vertically Written Japanese Text

2025-11-19 · Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara arxiv

Multimodal Large Language Models (MLLMs) have seen rapid advances in recent years and are now being applied to visual document understanding tasks. They are expected to process a wide range of document images across lang…

Harnessing PDF Data for Improving Japanese Large Multimodal Models

2025-02-20 · Jeonghun Baek, Akiko Aizawa, Kiyoharu Aizawa

Large Multimodal Models (LMMs) have demonstrated strong performance in English, but their effectiveness in Japanese remains limited due to the lack of high-quality training data. Current Japanese LMMs often rely on trans…

Optical Character Recognition (OCR)

A Corpus for English-Japanese Multimodal Neural Machine Translation with Comparable Sentences

2020-10-17 · Andrew Merritt, Chenhui Chu, Yuki Arase

Multimodal neural machine translation (NMT) has become an increasingly important area of research over the years because additional modalities, such as image data, can provide more context to textual data. Furthermore, t…

Image CaptioningMachine TranslationNMTSentence+1