paper-with-me

홈 › Papers

Crowdsourcing an OCR Gold Standard for a German and French Heritage Corpus

2016-05-01 · LREC 2016 5 · Simon Clematide, Lenz Furrer, Martin Volk

Crowdsourcing approaches for post-correction of OCR output (Optical Character Recognition) have been successfully applied to several historic text collections. We report on our crowd-correction platform Kokos, which we built to improve the OCR quality of the digitized yearbooks of the Swiss Alpine Club (SAC) from the 19th century. This multilingual heritage corpus consists of Alpine texts mainly written in German and French, all typeset in Antiqua font. Finding and engaging volunteers for correcting large amounts of pages into high quality text requires a carefully designed user interface, an easy-to-use workflow, and continuous efforts for keeping the participants motivated. More than 180,000 characters on about 21,000 pages were corrected by volunteers in about 7 month, achieving an OCR gold standard with a systematically evaluated accuracy of 99.7{\%} on the word level. The crowdsourced OCR gold standard and the corresponding original OCR recognition results from Abby FineReader 7 for each page are available as a resource. Additionally, the scanned images (300dpi) of all pages are included in order to facilitate tests with other OCR software.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Overview of the Second BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora

2017-08-01 · WS 2017 8 · Pierre Zweigenbaum, Serge Sharoff, Reinhard Rapp

This paper presents the BUCC 2017 shared task on parallel sentence extraction from comparable corpora. It recalls the design of the datasets, presents their final construction and statistics and the methods used to evalu…

Machine TranslationSentence

Strategies and Challenges for Crowdsourcing Regional Dialect Perception Data for Swiss German and Swiss French

2018-05-01 · LREC 2018 5 · Jean-Philippe Goldman, Simon Clematide, Mathieu Avanzi, T 외

Exposing ambiguities in a relation-extraction gold standard with crowdsourcing

2015-05-23 · Tong Shu Li, Benjamin M. Good, Andrew I. Su

Semantic relation extraction is one of the frontiers of biomedical natural language processing research. Gold standards are key tools for advancing this research. It is challenging to generate these standards because of …

RelationRelation Extraction

FOLK-Gold ― A Gold Standard for Part-of-Speech-Tagging of Spoken German

2016-05-01 · LREC 2016 5 · Swantje Westpfahl, Thomas Schmidt

In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers ― transcription (modified orthography), normalization (standard o…

LemmatizationPart-Of-Speech TaggingPOS

Supervised Rhyme Detection with Siamese Recurrent Networks

2018-08-01 · COLING 2018 8 · Thomas Haider, Jonas Kuhn

We present the first supervised approach to rhyme detection with Siamese Recurrent Networks (SRN) that offer near perfect performance (97{\%} accuracy) with a single model on rhyme pairs for German, English and French, a…

Binary ClassificationGeneral Classification