paper-with-me

Papers

Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents

2025-02-06 · Ilia Karmanov, Amala Sanjay Deshmukh, Lukas Voegtle, Philipp Fischer, Kateryna Chumachenko, Timo Roman, Jarno Seppänen, Jupinder Parmar, Joseph Jennings, Andrew Tao, Karan Sapra

Optical Character Recognition (OCR) technology is widely used to extract text from images of documents, facilitating efficient digitization and data retrieval. However, merely extracting text is insufficient when dealing with complex documents. Fully comprehending such documents requires an understanding of their structure -- including formatting, formulas, tables, and the reading order of multiple blocks and columns across multiple pages -- as well as semantic information for detecting elements like footnotes and image captions. This comprehensive understanding is crucial for downstream tasks such as retrieval, document question answering, and data curation for training Large Language Models (LLMs) and Vision Language Models (VLMs). To address this, we introduce \'Eclair, a general-purpose text-extraction tool specifically designed to process a wide range of document types. Given an image, \'Eclair is able to extract formatted text in reading order, along with bounding boxes and their corresponding semantic classes. To thoroughly evaluate these novel capabilities, we introduce our diverse human-annotated benchmark for document-level OCR and semantic classification. \'Eclair achieves state-of-the-art accuracy on this benchmark, outperforming other methods across key metrics. Additionally, we evaluate \'Eclair on established benchmarks, demonstrating its versatility and strength across several evaluation standards.

📄 PDF Abstract BibTeX arXiv:2502.04223

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningOptical Character RecognitionOptical Character Recognition (OCR)Question AnsweringRetrieval

Similar Papers 제목 키워드 기반

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs

2026-07-01 · Chahan Vidal-Gorène, Nadi Tomeh, Victoria Khurshudyan arxiv

This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We introduce a new annotated dataset of 66 pages and compare geometric h…

TRIE++: Towards End-to-End Information Extraction from Visually Rich Documents

2022-07-14 · Zhanzhan Cheng, Peng Zhang, Can Li, Qiao Liang 외

Recently, automatically extracting information from visually rich documents (e.g., tickets and resumes) has become a hot and vital research topic due to its widespread commercial value. Most existing methods divide this …

global-optimizationLanguage Modelling

LECTOR: Summarizing E-book Reading Content for Personalized Student Support

2025-05-12 · Erwin Daniel López Zapata, Cheng Tang, Valdemar Švábenský, Fumiya Okubo 외

Educational e-book platforms provide valuable information to teachers and researchers through two main sources: reading activity data and reading content data. While reading activity data is commonly used to analyze lear…

ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction

2026-01-09 · Tingwei Xie, Jinxin He, Yonghong Song arxiv

The efficacy of Multimodal Transformers in visually-rich document understanding (VrDU) is critically constrained by two inherent limitations: the lack of explicit modeling for logical reading order and the interference o…

Key Information Extraction

Combining Deep Learning and Reasoning for Address Detection in Unstructured Text Documents

2022-02-07 · AAAI Workshop CLeaR 2022 2 · Matthias Engelbach, Dennis Klau, Jens Drawehn, Maximilien Kintz

Extracting information from unstructured text documents is a demanding task, since these documents can have a broad variety of different layouts and a non-trivial reading order, like it is the case for multi-column docum…

Deep Learning