paper-with-me

홈 › Papers

Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents

2020-05-01 · LREC 2020 5 · Svetla Koeva, Nikola Obreshkov, Martin Yalamov

The paper presents the Bulgarian MARCELL corpus, part of a recently developed multilingual corpus representing the national legislation in seven European countries and the NLP pipeline that turns the web crawled data into structured, linguistically annotated dataset. The Bulgarian data is web crawled, extracted from the original HTML format, filtered by document type, tokenised, sentence split, tagged and lemmatised with a fine-grained version of the Bulgarian Language Processing Chain, dependency parsed with NLP- Cube, annotated with named entities (persons, locations, organisations and others), noun phrases, IATE terms and EuroVoc descriptors. An orchestrator process has been developed to control the NLP pipeline performing an end-to-end data processing and annotation starting from the documents identification and ending in the generation of statistical reports. The Bulgarian MARCELL corpus consists of 25,283 documents (at the beginning of November 2019), which are classified into eleven types.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Linguistic Analysis Processing Line for Bulgarian

2012-05-01 · LREC 2012 5 · Aleks Savkov, ar, Laska Laskova, Stanislava Kancheva 외

This paper presents a linguistic processing pipeline for Bulgarian including morphological analysis, lemmatization and syntactic analysis of Bulgarian texts. The morphological analysis is performed by three modules ― t…

Language ModellingLemmatizationMachine TranslationManagement+3

Annotation of Clinical Narratives in Bulgarian language

2017-09-01 · RANLP 2017 9 · Ivajlo Radev, Kiril Simov, Galia Angelova, Svetla Boytcheva

In this paper we describe annotation process of clinical texts with morphosyntactic and semantic information. The corpus contains 1,300 discharge letters in Bulgarian language for patients with Endocrinology and Metaboli…

ChunkingDependency ParsingInformation Retrieval

Comparative Analysis of Fine-tuned Deep Learning Language Models for ICD-10 Classification Task for Bulgarian Language

2021-09-01 · RANLP 2021 9 · Boris Velichkov, Sylvia Vassileva, Simeon Gerginov, Boris Kraychev 외

The task of automatic diagnosis encoding into standard medical classifications and ontologies, is of great importance in medicine - both to support the daily tasks of physicians in the preparation and reporting of clinic…

Categorisation of Bulgarian Legislative Documents

2020-09-01 · CLIB 2020 9 · Nikola Obreshkov, Martin Yalamov, Svetla Koeva

The paper presents the categorisation of Bulgarian MARCELL corpus in toplevel EuroVoc domains. The Bulgarian MARCELL corpus is part of a recently developed multilingual corpus representing the national legislation in sev…

Term Extraction

On Detecting Noun-Adjective Agreement Errors in Bulgarian Language Using GATE

2014-11-03 · Nadezhda Borisova, Grigor Iliev, Elena Karashtranova

In this article, we describe an approach for automatic detection of noun-adjective agreement errors in Bulgarian texts by explaining the necessary steps required to develop a simple Java-based language processing applica…

Retrieval