paper-with-me

홈 › Papers

LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR

2026-01-20 · Said Taghadouini, Adrien Cavaillès, Baptiste Aubertin arxiv

We present LightOnOCR-2-1B, a 1B-parameter end-to-end multilingual vision--language model that converts document images (e.g., PDFs) into clean, naturally ordered text without brittle OCR pipelines. Trained on a large-scale, high-quality distillation mix with strong coverage of scans, French documents, and scientific PDFs, LightOnOCR-2 achieves state-of-the-art results on OlmOCR-Bench while being 9$\times$ smaller and substantially faster than prior best-performing models. We further extend the output format to predict normalized bounding boxes for embedded images, introducing localization during pretraining via a resume strategy and refining it with RLVR using IoU-based rewards. Finally, we improve robustness with checkpoint averaging and task-arithmetic merging. We release model checkpoints under Apache 2.0, and publicly release the dataset and LightOnOCR-bbox-bench evaluation under their respective licenses.

📄 PDF Abstract BibTeX arXiv:2601.14251

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis

2026-06-28 · Avisha Dilhara, Nevidu Jayatilleke arxiv

Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing …

Document AI

mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs

2023-07-13 · Gregor Geigle, Abhay Jain, Radu Timofte, Goran Glavaš

Modular vision-language models (Vision-LLMs) align pretrained image encoders with (frozen) large language models (LLMs) and post-hoc condition LLMs to `understand' the image input. With the abundance of readily available…

Image Captioning

Generalizing Multimodal Pre-training into Multilingual via Language Acquisition

2022-05-29 · Liang Zhang, Anwen Hu, Qin Jin

English-based Vision-Language Pre-training (VLP) has achieved great success in various downstream tasks. Some efforts have been taken to generalize this success to non-English languages through Multilingual Vision-Langua…

Language AcquisitionRetrievalText RetrievalVideo-Text Retrieval

M-MiniGPT4: Multilingual VLLM Alignment via Translated Data

2026-03-31 · Seung Hun Han, Youssef Mohamed, Mohamed Elhoseiny arxiv

This paper presents a Multilingual Vision Large Language Model, named M-MiniGPT4. Our model exhibits strong vision-language understanding (VLU) capabilities across 11 languages. We utilize a mixture of native multilingua…

ICU: Conquering Language Barriers in Vision-and-Language Modeling by Dividing the Tasks into Image Captioning and Language Understanding

2023-10-19 · Guojun Wu

Most multilingual vision-and-language (V&L) research aims to accomplish multilingual and multimodal capabilities within one model. However, the scarcity of multilingual captions for images has hindered the development. T…

Image CaptioningLanguage ModelingLanguage Modelling