paper-with-me

홈 › Papers

Bulgarian X-language Parallel Corpus

2012-05-01 · LREC 2012 5 · Svetla Koeva, Ivelina Stoyanova, Rositsa Dekova, Borislav Rizov, Angel Genov

The paper presents the methodology and the outcome of the compilation and the processing of the Bulgarian X-language Parallel Corpus (Bul-X-Cor) which was integrated as part of the Bulgarian National Corpus (BulNC). We focus on building representative parallel corpora which include a diversity of domains and genres, reflect the relations between Bulgarian and other languages and are consistent in terms of compilation methodology, text representation, metadata description and annotation conventions. The approaches implemented in the construction of Bul-X-Cor include using readily available text collections on the web, manual compilation (by means of Internet browsing) and preferably automatic compilation (by means of web crawling ― general and focused). Certain levels of annotation applied to Bul-X-Cor are taken as obligatory (sentence segmentation and sentence alignment), while others depend on the availability of tools for a particular language (morpho-syntactic tagging, lemmatisation, syntactic parsing, named entity recognition, word sense disambiguation, etc.) or for a particular task (word and clause alignment). To achieve uniformity of the annotation we have either annotated raw data from scratch or transformed the already existing annotation to follow the conventions accepted for BulNC. Finally, actual uses of the corpora are presented and conclusions are drawn with respect to future work.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)SentenceSentence segmentationWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Some Notes on p(e)re-Reduplication in Bulgarian and Ukrainian: A Corpus-based Study

2022-09-01 · CLIB 2022 9 · Ivan Derzhanski, Olena Siruk

We present a comparative study of p(e)re-reduplication in Bulgarian and Ukrainian, based on material from a parallel corpus of bilingual texts. We analyse all occurrences found in the corpus of close sequences and conjun…

Expanding Parallel Resources for Medium-Density Languages for Free

2012-05-01 · LREC 2012 5 · Georgi Iliev, Angel Genov

We discuss a previously proposed method for augmenting parallel corpora of limited size for the purposes of machine translation through monolingual paraphrasing of the source language. We develop a three-stage shallow pa…

Machine TranslationMorphological AnalysisTranslation

A Parallel WordNet for English, Swedish and Bulgarian

2020-05-01 · LREC 2020 5 · Krasimir Angelov

We present the parallel creation of a WordNet resource for Swedish and Bulgarian which is tightly aligned with the Princeton WordNet. The alignment is not only on the synset level, but also on word level, by matching wor…

ArticlesMachine TranslationText GenerationTranslation

A Bilingual Lexicosemantic Network of Bread Based on a Parallel Corpus

2020-09-01 · CLIB 2020 9 · Ivan Derzhanski, Olena Siruk

We present an experiment in using a corpus of Bulgarian and Ukrainian parallel texts for the automatised construction of a bilingual lexicosemantic network representing the semantic field of BREAD. We discuss the extract…

QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages

2016-05-01 · LREC 2016 5 · Arantxa Otegi, Nora Aranberri, Antonio Branco, Jan Haji{\v{c}} 외

This work presents parallel corpora automatically annotated with several NLP tools, including lemma and part-of-speech tagging, named-entity recognition and classification, named-entity disambiguation, word-sense disambi…

Cross-Lingual TransferEntity DisambiguationGeneral ClassificationLEMMA+7