paper-with-me

Papers

Harnessing PDF Data for Improving Japanese Large Multimodal Models

2025-02-20 · Jeonghun Baek, Akiko Aizawa, Kiyoharu Aizawa

Large Multimodal Models (LMMs) have demonstrated strong performance in English, but their effectiveness in Japanese remains limited due to the lack of high-quality training data. Current Japanese LMMs often rely on translated English datasets, restricting their ability to capture Japan-specific cultural knowledge. To address this, we explore the potential of Japanese PDF data as a training resource, an area that remains largely underutilized. We introduce a fully automated pipeline that leverages pretrained models to extract image-text pairs from PDFs through layout analysis, OCR, and vision-language pairing, removing the need for manual annotation. Additionally, we construct instruction data from extracted image-text pairs to enrich the training data. To evaluate the effectiveness of PDF-derived data, we train Japanese LMMs and assess their performance on the Japanese LMM Benchmark. Our results demonstrate substantial improvements, with performance gains ranging from 3.9% to 13.8% on Heron-Bench. Further analysis highlights the impact of PDF-derived data on various factors, such as model size and language models, reinforcing its value as a multimodal resource for Japanese LMMs. We plan to make the source code and data publicly available upon acceptance.

📄 PDF Abstract BibTeX arXiv:2502.14778

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

JDocQA: Japanese Document Question Answering Dataset for Generative Language Models

2024-03-28 · Eri Onami, Shuhei Kurita, Taiki Miyanishi, Taro Watanabe

Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites, and it is a truly demanding task as paper and electronic forms of documents are so common i…

HallucinationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Evaluating Multimodal Large Language Models on Vertically Written Japanese Text

2025-11-19 · Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara arxiv

Multimodal Large Language Models (MLLMs) have seen rapid advances in recent years and are now being applied to visual document understanding tasks. They are expected to process a wide range of document images across lang…

Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model

2024-10-30 · Keito Sasagawa, Koki Maeda, Issa Sugiura, Shuhei Kurita 외

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abun…

Language ModelingLanguage Modelling

JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

2024-10-22 · Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira 외

Accelerating research on Large Multimodal Models (LMMs) in non-English languages is crucial for enhancing user experiences across broader populations. In this paper, we introduce JMMMU (Japanese MMMU), the first large-sc…

Math

Evolutionary Optimization of Model Merging Recipes

2024-03-19 · Takuya Akiba, Makoto Shing, Yujin Tang, Qi Sun 외

Large language models (LLMs) have become increasingly capable, but their development often requires substantial computational resources. While model merging has emerged as a cost-effective promising approach for creating…

Evolutionary AlgorithmsMathmodel