paper-with-me

Papers

DriveThru: a Document Extraction Platform and Benchmark Datasets for Indonesian Local Language Archives

2024-11-14 · Mohammad Rifqi Farhansyah, Muhammad Zuhdi Fikri Johari, Afinzaki Amiral, Ayu Purwarianti, Kumara Ari Yuana, Derry Tanti Wijaya

Indonesia is one of the most diverse countries linguistically. However, despite this linguistic diversity, Indonesian languages remain underrepresented in Natural Language Processing (NLP) research and technologies. In the past two years, several efforts have been conducted to construct NLP resources for Indonesian languages. However, most of these efforts have been focused on creating manual resources thus difficult to scale to more languages. Although many Indonesian languages do not have a web presence, locally there are resources that document these languages well in printed forms such as books, magazines, and newspapers. Digitizing these existing resources will enable scaling of Indonesian language resource construction to many more languages. In this paper, we propose an alternative method of creating datasets by digitizing documents, which have not previously been used to build digital language resources in Indonesia. DriveThru is a platform for extracting document content utilizing Optical Character Recognition (OCR) techniques in its system to provide language resource building with less manual effort and cost. This paper also studies the utility of current state-of-the-art LLM for post-OCR correction to show the capability of increasing the character accuracy rate (CAR) and word accuracy rate (WAR) compared to off-the-shelf OCR.

📄 PDF Abstract BibTeX arXiv:2411.09318

Code (1)

ragambahasa/OCR-Correction 공식 구현

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

BibRank: Automatic Keyphrase Extraction Platform Using~Metadata

2023-10-13 · Abdelrhman Eldallal, Eduard Barbu

Automatic Keyphrase Extraction involves identifying essential phrases in a document. These keyphrases are crucial in various tasks such as document classification, clustering, recommendation, indexing, searching, summari…

ClusteringDocument ClassificationKeyphrase ExtractionText Simplification

Using Neighborhood Context to Improve Information Extraction from Visual Documents Captured on Mobile Phones

2021-08-23 · Kalpa Gunaratna, Vijay Srinivasan, Sandeep Nama, Hongxia Jin

Information Extraction from visual documents enables convenient and intelligent assistance to end users. We present a Neighborhood-based Information Extraction (NIE) approach that uses contextual language models and pays…

x.ent: R Package for Entities and Relations Extraction based on Unsupervised Learning and Document Structure

2015-04-23 · Nicolas Turenne, Tien Phan

Relation extraction with accurate precision is still a challenge when processing full text databases. We propose an approach based on cooccurrence analysis in each document for which we used document organization to impr…

EpidemiologyRelationRelation Extraction

Business Document Information Extraction: Towards Practical Benchmarks

2022-06-20 · Matyáš Skalický, Štěpán Šimsa, Michal Uřičář, Milan Šulc

Information extraction from semi-structured documents is crucial for frictionless business-to-business (B2B) communication. While machine learning problems related to Document Information Extraction (IE) have been studie…

ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images

2026-02-12 · Mathieu Sibue, Andres Muñoz Garza, Samuel Mensah, Pranav Shetty 외 arxiv

Enterprise documents, such as forms and reports, embed critical information for downstream applications like data archiving, automated workflows, and analytics. Although generalist Vision Language Models (VLMs) perform w…

Visual Question AnsweringInformation ExtractionRelation Extraction