Corpora and Processing Tools for Non-standard Contemporary and Diachronic Balkan Slavic
The paper describes three corpora of different varieties of BS that are currently being developed with the goal of providing data for the analysis of the diatopic and diachronic variation in non-standard Balkan Slavic. The corpora includes spoken materials from Torlak, Macedonian dialects, as well as the manuscripts of pre-standardized Bulgarian. Apart from the texts, tools for PoS annotation and lemmatization for all varieties are being created, as well as syntactic parsing for Torlak and Bulgarian varieties. The corpora are built using a unified methodology, relying on the pest practices and state-of-the-art methods from the field. The uniform methodology allows the contrastive analysis of the data from different varieties. The corpora under construction can be considered a crucial contribution to the linguistic research on the languages in the Balkans as they provide the lacking data needed for the studies of linguistic variation in the Balkan Slavic, and enable the comparison of the said varieties with other neighbouring languages.
Code (0)
등록된 구현이 없습니다.
Tasks
LemmatizationPOSSimilar Papers 제목 키워드 기반
A Diachronic Corpus for Romanian (RoDia)
This paper describes a Romanian Dependency Treebank, built at the Al. I. Cuza University (UAIC), and a special OCR techniques used to build it. The corpus has rich morphological and syntactic annotation. There are few an…
Information RetrievalOptical Character Recognition (OCR)Question AnsweringStandardizing linguistic data: method and tools for annotating (pre-orthographic) French
With the development of big corpora of various periods, it becomes crucial to standardise linguistic annotation (e.g. lemmas, POS tags, morphological annotation) to increase the interoperability of the data produced, des…
POSDUKweb: Diachronic word representations from the UK Web Archive corpus
Lexical semantic change (detecting shifts in the meaning and usage of words) is an important task for social and cultural studies as well as for Natural Language Processing applications. Diachronic word embeddings (time-…
Change DetectionDiachronic Word EmbeddingsWord EmbeddingsTools for Building a Corpus to Study the Historical and Geographical Variation of the Romanian Language
Contemporary standard language corpora are ideal for NLP. There are few morphologically and syntactically annotated corpora for Romanian, and those existing or in progress only deal with the Contemporary Romanian standar…
Setting Up Bilingual Comparable Corpora with Non-Contemporary Languages
This paper presents the project “Les corpora latins et français: une fabrique pour l’accès à la représentation des connaissances” (Latin and French Corpora: a Factory For Accessing Knowledge Representation) whose focus i…