paper-with-me

홈 › Papers

Collection and Annotation of the Romanian Legal Corpus

2020-05-01 · LREC 2020 5 · Dan Tufi{\textcommabelow{s}}, Maria Mitrofan, Vasile P{\u{a}}i{\textcommabelow{s}}, Radu Ion, Andrei Coman

We present the Romanian legislative corpus which is a valuable linguistic asset for the development of machine translation systems, especially for under-resourced languages. The knowledge that can be extracted from this resource is necessary for a deeper understanding of how law terminology is used and how it can be made more consistent. At this moment the corpus contains more than 140k documents representing the legislative body of Romania. This corpus is processed and annotated at different levels: linguistically (tokenized, lemmatized and pos-tagged), dependency parsed, chunked, named entities identified and labeled with IATE terms and EUROVOC descriptors. Each annotated document has a CONLL-U Plus format consisting in 14 columns, in addition to the standard 10-column format, four other types of annotations were added. Moreover the repository will be periodically updated as new legislative texts are published. These will be automatically collected and transmitted to the processing and annotation pipeline. The access to the corpus will be done through ELRC infrastructure.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationPOSTranslation

Similar Papers 제목 키워드 기반

ROMBAC: The Romanian Balanced Annotated Corpus

2012-05-01 · LREC 2012 5 · Radu Ion, Elena Irimia, Dan {\c{S}}tef{\u{a}}nescu, Dan Tufi{\textcommabelow{s}}

This article describes the collecting, processing and validation of a large balanced corpus for Romanian. The annotation types and structure of the corpus are briefly reviewed. It was constructed at the Research Institut…

ChunkingLemmatizationPOSPOS Tagging

Identifying Draft Bills Impacting Existing Legislation: a Case Study on Romanian

2022-06-01 · LREC 2022 6 · Corina Ceausu, Sergiu Nisioi

In our paper, we present a novel corpus of historical legal documents on the Romanian public procurement legislation and an annotated subset of draft bills that have been screened by legal experts and identified as impac…

Romanian micro-blogging named entity recognition including health-related entities

2022-10-01 · SMM4H (COLING) 2022 10 · Vasile Pais, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan 외

This paper introduces a manually annotated dataset for named entity recognition (NER) in micro-blogging text for Romanian language. It contains gold annotations for 9 entity classes and expressions: persons, locations, o…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Named Entity Recognition in the Romanian Legal Domain

2021-11-01 · EMNLP (NLLP) 2021 11 · Vasile Pais, Maria Mitrofan, Carol Luca Gasan, Vlad Coneschi 외

Recognition of named entities present in text is an important step towards information extraction and natural language understanding. This work presents a named entity recognition system for the Romanian legal domain. Th…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Understanding+1

CoRoLa --- The Reference Corpus of Contemporary Romanian Language

2014-05-01 · LREC 2014 5 · Verginica Barbu Mititelu, Elena Irimia, Dan Tufi{\textcommabelow{s}}

We present the project of creating CoRoLa, a reference corpus of contemporary Romanian (from 1945 onwards). In the international context, the project finds its place among the initiatives of gathering huge collections of…

LemmatizationSentence