paper-with-me

홈 › Papers

The IPR-cleared Corpus of Contemporary Written and Spoken Romanian Language

2016-05-01 · LREC 2016 5 · Dan Tufi{\textcommabelow{s}}, Verginica Barbu Mititelu, Elena Irimia, {\textcommabelow{S}}tefan Daniel Dumitrescu, Tiberiu Boro{\textcommabelow{s}}

The article describes the current status of a large national project, CoRoLa, aiming at building a reference corpus for the contemporary Romanian language. Unlike many other national corpora, CoRoLa contains only - IPR cleared texts and speech data, obtained from some of the country{'}s most representative publishing houses, broadcasting agencies, editorial offices, newspapers and popular bloggers. For the written component 500 million tokens are targeted and for the oral one 300 hours of recordings. The choice of texts is done according to their functional style, domain and subdomain, also with an eye to the international practice. A metadata file (following the CMDI model) is associated to each text file. Collected texts are cleaned and transformed in a format compatible with the tools for automatic processing (segmentation, tokenization, lemmatization, part-of-speech tagging). The paper also presents up-to-date statistics about the structure of the corpus almost two years before its official launching. The corpus will be freely available for searching. Users will be able to download the results of their searches and those original files when not against stipulations in the protocols we have with text providers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

LemmatizationPart-Of-Speech Tagging

Similar Papers 제목 키워드 기반

A Diachronic Corpus for Romanian (RoDia)

2017-09-01 · RANLP 2017 9 · Ludmila Malahov, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Alex Colesnicov, ru

This paper describes a Romanian Dependency Treebank, built at the Al. I. Cuza University (UAIC), and a special OCR techniques used to build it. The corpus has rich morphological and syntactic annotation. There are few an…

Information RetrievalOptical Character Recognition (OCR)Question Answering

Tools for Building a Corpus to Study the Historical and Geographical Variation of the Romanian Language

2017-09-01 · RANLP 2017 9 · Victoria Bobicev, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Cenel Augusto Perez

Contemporary standard language corpora are ideal for NLP. There are few morphologically and syntactically annotated corpora for Romanian, and those existing or in progress only deal with the Contemporary Romanian standar…

Resources in Underrepresented Languages: Building a Representative Romanian Corpus

2020-05-01 · LREC 2020 5 · Ludmila Midrigan - Ciochina, Victoria Boyd, Lucila Sanchez-Ortega, Diana Malancea{\_}Malac 외

The effort in the field of Linguistics to develop theories that aim to explain language-dependent effects on language processing is greatly facilitated by the availability of reliable resources representing different lan…

CoRoLa --- The Reference Corpus of Contemporary Romanian Language

2014-05-01 · LREC 2014 5 · Verginica Barbu Mititelu, Elena Irimia, Dan Tufi{\textcommabelow{s}}

We present the project of creating CoRoLa, a reference corpus of contemporary Romanian (from 1945 onwards). In the international context, the project finds its place among the initiatives of gathering huge collections of…

LemmatizationSentence

SYN2015: Representative Corpus of Contemporary Written Czech

2016-05-01 · LREC 2016 5 · Michal K{\v{r}}en, V{\'a}clav Cvr{\v{c}}ek, Tom{\'a}{\v{s}} {\v{C}}apka, Anna {\v{C}}erm{\'a}kov{\'a} 외

The paper concentrates on the design, composition and annotation of SYN2015, a new 100-million representative corpus of contemporary written Czech. SYN2015 is a sequel of the representative corpora of the SYN series that…

text-classificationText Classification