paper-with-me

홈 › Papers

CAMIO: A Corpus for OCR in Multiple Languages

2022-06-01 · LREC 2022 6 · Michael Arrigo, Stephanie Strassel, Nolan King, Thao Tran, Lisa Mason

CAMIO (Corpus of Annotated Multilingual Images for OCR) is a new corpus created by Linguistic Data Consortium to serve as a resource to support the development and evaluation of optical character recognition (OCR) and related technologies for 35 languages across 24 unique scripts. The corpus comprises nearly 70,000 images of machine printed text, covering a wide variety of topics and styles, document domains, attributes and scanning/capture artifacts. Most images have been exhaustively annotated for text localization, resulting in over 2.3M line-level bounding boxes. For 13 of the 35 languages, 1250 images/language have been further annotated with orthographic transcriptions of each line plus specification of reading order, yielding over 2.4M tokens of transcribed text. The resulting annotations are represented in a comprehensive XML output format defined for this corpus. The paper discusses corpus design and implementation, challenges encountered, baseline performance results obtained on the corpus for text localization and OCR decoding, and plans for corpus publication.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction

2018-08-01 · COLING 2018 8 · Nurendra Choudhary, Rajat Singh, Vijjini Anvesh Rao, Manish Shrivastava

In this paper, we leverage social media platforms such as twitter for developing corpus across multiple languages. The corpus creation methodology is applicable for resource-scarce languages provided the speakers of that…

Sentiment Analysis

AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages

2020-04-30 · Anoop Kunchukuttan, Divyanshu Kakwani, Satish Golla, Gokul N. C. 외

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We crea…

Word Embeddings

The Multilingual TEDx Corpus for Speech Recognition and Translation

2021-02-02 · Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman, Roldano Cattoni 외

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx t…

speech-recognitionSpeech RecognitionTranslation

The NewSoMe Corpus: A Unifying Opinion Annotation Framework across Genres and in Multiple Languages

2014-05-01 · LREC 2014 5 · Roser Saur{\'\i}, Judith Domingo, Toni Badia

We present the NewSoMe (News and Social Media) Corpus, a set of subcorpora with annotations on opinion expressions across genres (news reports, blogs, product reviews and tweets) and covering multiple languages (English,…

Information RetrievalOpinion Mining

ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

2024-03-27 · Injy Hamed, Fadhl Eryani, David Palfreyman, Nizar Habash

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)LemmatizationPart-Of-Speech Tagging+2