paper-with-me

Papers

Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models

2026-04-02 · Issa Sugiura, Keito Sasagawa, Keisuke Nakao, Koki Maeda, Ziqi Yin, Zhishen Yang, Shuhei Kurita, Yusuke Oda, Ryoko Tokuhisa, Daisuke Kawahara, Naoaki Okazaki arxiv

Developing vision-language models (VLMs) that generalize across diverse tasks requires large-scale training datasets with diverse content. In English, such datasets are typically constructed by aggregating and curating numerous existing visual question answering (VQA) resources. However, this strategy does not readily extend to other languages, where VQA datasets remain limited in both scale and domain coverage, posing a major obstacle to building high-quality multilingual and non-English VLMs. In this work, we introduce Jagle, the largest Japanese multimodal post-training dataset to date, comprising approximately 9.2 million instances across diverse tasks. Rather than relying on existing VQA datasets, we collect heterogeneous source data, including images, image-text pairs, and PDF documents, and generate VQA pairs through multiple strategies such as VLM-based QA generation, translation, and text rendering. Experiments demonstrate that a 2.2B model trained with Jagle achieves strong performance on Japanese tasks, surpassing InternVL3.5-2B in average score across ten Japanese evaluation tasks and approaching within five points of Qwen3-VL-2B-Instruct. Furthermore, combining Jagle with FineVision does not degrade English performance; instead, it improves English performance compared to training with FineVision alone. To facilitate reproducibility and future research, we release the dataset, trained models, and code.

📄 PDF Abstract BibTeX arXiv:2604.02048

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

2024-10-22 · Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira 외

Accelerating research on Large Multimodal Models (LMMs) in non-English languages is crucial for enhancing user experiences across broader populations. In this paper, we introduce JMMMU (Japanese MMMU), the first large-sc…

Math

DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering

2025-11-30 · Toshiki Katsube, Taiga Fukuhara, Kenichiro Ando, Yusuke Mukuta 외 arxiv

This work addresses the scarcity of high-quality, large-scale resources for Japanese Vision-and-Language (V&L) modeling. We present a scalable and reproducible pipeline that integrates large-scale web collection with rig…

Visual Question AnsweringImage Captioning

Harnessing PDF Data for Improving Japanese Large Multimodal Models

2025-02-20 · Jeonghun Baek, Akiko Aizawa, Kiyoharu Aizawa

Large Multimodal Models (LMMs) have demonstrated strong performance in English, but their effectiveness in Japanese remains limited due to the lack of high-quality training data. Current Japanese LMMs often rely on trans…

Optical Character Recognition (OCR)

Empirical Analysis of Training Strategies of Transformer-based Japanese Chit-chat Systems

2021-09-11 · Hiroaki Sugiyama, Masahiro Mizukami, Tsunehiro Arimoto, Hiromi Narimatsu 외

In recent years, several high-performance conversational systems have been proposed based on the Transformer encoder-decoder model. Although previous studies analyzed the effects of the model parameters and the decoding …

Decoder

Empirical Analysis of Training Strategies of Transformer-based Japanese Chit-chat Systems

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In recent years, several high-performance conversational systems have been proposed based on the Transformer encoder-decoder model. Although previous studies analyzed the effects of the model parameters and the decoding …

Decoder