paper-with-me

홈 › Papers

ROMBAC: The Romanian Balanced Annotated Corpus

2012-05-01 · LREC 2012 5 · Radu Ion, Elena Irimia, Dan {\c{S}}tef{\u{a}}nescu, Dan Tufi{\textcommabelow{s}}

This article describes the collecting, processing and validation of a large balanced corpus for Romanian. The annotation types and structure of the corpus are briefly reviewed. It was constructed at the Research Institute for Artificial Intelligence of the Romanian Academy in the context of an international project (METANET4U). The processing covers tokenization, POS-tagging, lemmatization and chunking. The corpus is in XML format generated by our in-house annotation tools; the corpus encoding schema is XCES compliant and the metadata specification is conformant to the METANET recommendations. To the best of our knowledge, this is the first large and richly annotated corpus for Romanian. ROMBAC is intended to be the foundation of a linguistic environment containing a reference corpus for contemporary Romanian and a comprehensive collection of interoperable processing tools.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ChunkingLemmatizationPOSPOS Tagging

Similar Papers 제목 키워드 기반

Automatic Extraction of the Romanian Academic Word List: Data and Methods

2023-07-29 · Ana-Maria Bucur, Andreea Dincă, Mădălina Chitez, Roxana Rogobete

This paper presents the methodology and data used for the automatic extraction of the Romanian Academic Word List (Ro-AWL). Academic Word Lists are useful in both L2 and L1 teaching contexts. For the Romanian language, n…

POS

Tools for Building a Corpus to Study the Historical and Geographical Variation of the Romanian Language

2017-09-01 · RANLP 2017 9 · Victoria Bobicev, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Cenel Augusto Perez

Contemporary standard language corpora are ideal for NLP. There are few morphologically and syntactically annotated corpora for Romanian, and those existing or in progress only deal with the Contemporary Romanian standar…

Romanian TimeBank: An Annotated Parallel Corpus for Temporal Information

2012-05-01 · LREC 2012 5 · Corina For{\u{a}}scu, Dan Tufi{\c{s}}

The paper describes the main steps for the construction, annotation and validation of the Romanian version of the TimeBank corpus. Starting from the English TimeBank corpus ― the reference annotated corpus in the tempo…

Information RetrievalMachine TranslationQuestion AnsweringTAG

The Romanian Corpus Annotated with Verbal Multiword Expressions

2019-08-01 · WS 2019 8 · Verginica Barbu Mititelu, Mihaela Cristescu, Mihaela Onofrei

This paper reports on the Romanian journalistic corpus annotated with verbal multiword expressions following the PARSEME guidelines. The corpus is sentence split, tokenized, part-of-speech tagged, lemmatized, syntactical…

DiversitySentence

Introducing RONEC -- the Romanian Named Entity Corpus

2019-09-03 · Stefan Daniel Dumitrescu, Andrei-Marius Avram

We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in ~5000 annotated sentences, belonging to 16 distinct classes. The sentences have been extracted from a copy-…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)