paper-with-me

Papers

Corpora and Processing Tools for Non-standard Contemporary and Diachronic Balkan Slavic

2019-09-01 · RANLP 2019 9 · Teodora Vukovic, Nora Muheim, Olivier Winist{\"o}rfer, Ivan {\v{S}}imko, Anastasia Makarova, Sanja Bradjan

The paper describes three corpora of different varieties of BS that are currently being developed with the goal of providing data for the analysis of the diatopic and diachronic variation in non-standard Balkan Slavic. The corpora includes spoken materials from Torlak, Macedonian dialects, as well as the manuscripts of pre-standardized Bulgarian. Apart from the texts, tools for PoS annotation and lemmatization for all varieties are being created, as well as syntactic parsing for Torlak and Bulgarian varieties. The corpora are built using a unified methodology, relying on the pest practices and state-of-the-art methods from the field. The uniform methodology allows the contrastive analysis of the data from different varieties. The corpora under construction can be considered a crucial contribution to the linguistic research on the languages in the Balkans as they provide the lacking data needed for the studies of linguistic variation in the Balkan Slavic, and enable the comparison of the said varieties with other neighbouring languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

LemmatizationPOS

Similar Papers 제목 키워드 기반

A Diachronic Corpus for Romanian (RoDia)

2017-09-01 · RANLP 2017 9 · Ludmila Malahov, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Alex Colesnicov, ru

This paper describes a Romanian Dependency Treebank, built at the Al. I. Cuza University (UAIC), and a special OCR techniques used to build it. The corpus has rich morphological and syntactic annotation. There are few an…

Information RetrievalOptical Character Recognition (OCR)Question Answering

Standardizing linguistic data: method and tools for annotating (pre-orthographic) French

2020-11-22 · Simon Gabay, Thibault Clérice, Jean-Baptiste Camps, Jean-Baptiste Tanguy 외

With the development of big corpora of various periods, it becomes crucial to standardise linguistic annotation (e.g. lemmas, POS tags, morphological annotation) to increase the interoperability of the data produced, des…

POS

DUKweb: Diachronic word representations from the UK Web Archive corpus

2021-07-02 · Adam Tsakalidis, Pierpaolo Basile, Marya Bazzi, Mihai Cucuringu 외

Lexical semantic change (detecting shifts in the meaning and usage of words) is an important task for social and cultural studies as well as for Natural Language Processing applications. Diachronic word embeddings (time-…

Change DetectionDiachronic Word EmbeddingsWord Embeddings

Tools for Building a Corpus to Study the Historical and Geographical Variation of the Romanian Language

2017-09-01 · RANLP 2017 9 · Victoria Bobicev, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Cenel Augusto Perez

Contemporary standard language corpora are ideal for NLP. There are few morphologically and syntactically annotated corpora for Romanian, and those existing or in progress only deal with the Contemporary Romanian standar…

Setting Up Bilingual Comparable Corpora with Non-Contemporary Languages

2022-06-01 · LREC (BUCC) 2022 6 · Helena Bermudez Sabel, Francesca Dell’Oro, Cyrielle Montrichard, Corinne Rossari

This paper presents the project “Les corpora latins et français: une fabrique pour l’accès à la représentation des connaissances” (Latin and French Corpora: a Factory For Accessing Knowledge Representation) whose focus i…