paper-with-me

홈 › Papers

MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding

2025-11-13 · Ketong Chen, Yuhao Chen, Yang Xue arxiv

Despite the rapid progress of Vision-Language Models (VLMs), their capabilities are inadequately assessed by existing benchmarks, which are predominantly English-centric, feature simplistic layouts, and support limited tasks. Consequently, they fail to evaluate model performance for Visually Rich Document Understanding (VRDU), a critical challenge involving complex layouts and dense text. To address this, we introduce DocWeaver, a novel multi-agent pipeline that leverages Large Language Models to automatically generate a new benchmark. The result is MosaicDoc, a large-scale, bilingual (Chinese and English) resource designed to push the boundaries of VRDU. Sourced from newspapers and magazines, MosaicDoc features diverse and complex layouts (including multi-column and non-Manhattan), rich stylistic variety from 196 publishers, and comprehensive multi-task annotations (OCR, VQA, reading order, and localization). With 72K images and over 600K QA pairs, MosaicDoc serves as a definitive benchmark for the field. Our extensive evaluation of state-of-the-art models on this benchmark reveals their current limitations in handling real-world document complexity and charts a clear path for future research.

📄 PDF Abstract BibTeX arXiv:2511.09919

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The mutual exclusivity bias of bilingual visually grounded speech models

2025-06-04 · Dan Oneata, Leanne Nortje, Yevgen Matusevych, Herman Kamper

Mutual exclusivity (ME) is a strategy where a novel word is associated with a novel object rather than a familiar one, facilitating language learning in children. Recent work has found an ME bias in a visually grounded s…

Movie101v2: Improved Movie Narration Benchmark

2024-04-20 · Zihao Yue, Yepeng Zhang, Ziheng Wang, Qin Jin

Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences. Unlike standard video captioning, it involves not only describing key visual details but also inferring pl…

Video Captioning

Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples

2023-03-30 · Hyeonggon Ryu, Arda Senocak, In So Kweon, Joon Son Chung

The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equal number of spoken captions from both l…

Cross-Modal RetrievalRetrieval

Models of Visually Grounded Speech Signal Pay Attention To Nouns: a Bilingual Experiment on English and Japanese

2019-02-08 · William N. Havard, Jean-Pierre Chevrot, Laurent Besacier

We investigate the behaviour of attention in neural models of visually grounded speech trained on two languages: English and Japanese. Experimental results show that attention focuses on nouns and this behaviour holds tr…

Retrieval

Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding

2025-11-01 · Haneen Al-Homoud, Asma Ibrahim, Murtadha Al-Jubran, Fahad Al-Otaibi 외 arxiv

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 milli…