Bulgarian X-language Parallel Corpus
The paper presents the methodology and the outcome of the compilation and the processing of the Bulgarian X-language Parallel Corpus (Bul-X-Cor) which was integrated as part of the Bulgarian National Corpus (BulNC). We focus on building representative parallel corpora which include a diversity of domains and genres, reflect the relations between Bulgarian and other languages and are consistent in terms of compilation methodology, text representation, metadata description and annotation conventions. The approaches implemented in the construction of Bul-X-Cor include using readily available text collections on the web, manual compilation (by means of Internet browsing) and preferably automatic compilation (by means of web crawling ― general and focused). Certain levels of annotation applied to Bul-X-Cor are taken as obligatory (sentence segmentation and sentence alignment), while others depend on the availability of tools for a particular language (morpho-syntactic tagging, lemmatisation, syntactic parsing, named entity recognition, word sense disambiguation, etc.) or for a particular task (word and clause alignment). To achieve uniformity of the annotation we have either annotated raw data from scratch or transformed the already existing annotation to follow the conventions accepted for BulNC. Finally, actual uses of the corpora are presented and conclusions are drawn with respect to future work.
Code (0)
등록된 구현이 없습니다.
Tasks
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)SentenceSentence segmentationWord Sense DisambiguationSimilar Papers 제목 키워드 기반
Some Notes on p(e)re-Reduplication in Bulgarian and Ukrainian: A Corpus-based Study
We present a comparative study of p(e)re-reduplication in Bulgarian and Ukrainian, based on material from a parallel corpus of bilingual texts. We analyse all occurrences found in the corpus of close sequences and conjun…
Expanding Parallel Resources for Medium-Density Languages for Free
We discuss a previously proposed method for augmenting parallel corpora of limited size for the purposes of machine translation through monolingual paraphrasing of the source language. We develop a three-stage shallow pa…
Machine TranslationMorphological AnalysisTranslationA Parallel WordNet for English, Swedish and Bulgarian
We present the parallel creation of a WordNet resource for Swedish and Bulgarian which is tightly aligned with the Princeton WordNet. The alignment is not only on the synset level, but also on word level, by matching wor…
ArticlesMachine TranslationText GenerationTranslationA Bilingual Lexicosemantic Network of Bread Based on a Parallel Corpus
We present an experiment in using a corpus of Bulgarian and Ukrainian parallel texts for the automatised construction of a bilingual lexicosemantic network representing the semantic field of BREAD. We discuss the extract…
QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages
This work presents parallel corpora automatically annotated with several NLP tools, including lemma and part-of-speech tagging, named-entity recognition and classification, named-entity disambiguation, word-sense disambi…
Cross-Lingual TransferEntity DisambiguationGeneral ClassificationLEMMA+7