paper-with-me

홈 › Papers

Recycling and Comparing Morphological Annotation Models for Armenian Diachronic-Variational Corpus Processing

2020-12-01 · VarDial (COLING) 2020 12 · Chahan Vidal-Gorène, Victoria Khurshudyan, Anaïd Donabédian-Demopoulos

Armenian is a language with significant variation and unevenly distributed NLP resources for different varieties. An attempt is made to process an RNN model for morphological annotation on the basis of different Armenian data (provided or not with morphologically annotated corpora), and to compare the annotation results of RNN and rule-based models. Different tests were carried out to evaluate the reuse of an unspecialized model of lemmatization and POS-tagging for under-resourced language varieties. The research focused on three dialects and further extended to Western Armenian with a mean accuracy of 94,00 % in lemmatization and 97,02% in POS-tagging, as well as a possible reusability of models to cover different other Armenian varieties. Interestingly, the comparison of an RNN model trained on Eastern Armenian with the Eastern Armenian National Corpus rule-based model applied to Western Armenian showed an enhancement of 19% in parsing. This model covers 88,79% of a short heterogeneous dataset in Western Armenian, and could be a baseline for a massive corpus annotation in that standard. It is argued that an RNN-based model can be a valid alternative to a rule-based one giving consideration to such factors as time-consumption, reusability for different varieties of a target language and significant qualitative results in morphological annotation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

LemmatizationPOSPOS Taggingvalid

Similar Papers 제목 키워드 기반

Eastern Armenian National Corpus: State of the Art and Perspectives

2022-06-01 · DigitAm (LREC) 2022 6 · Victoria Khurshudyan, Timofey Arkhangelskiy, Misha Daniel, Vladimir Plungian 외

Eastern Armenian National Corpus (EANC) is a comprehensive corpus of Modern Eastern Armenian with about 110 million tokens, covering written and oral discourses from the mid-19th century to the present. The corpus is pro…

A Free/Open-Source Morphological Transducer for Western Armenian

2022-06-01 · DigitAm (LREC) 2022 6 · Hossep Dolatian, Daniel Swanson, Jonathan Washington

We present a free/open-source morphological transducer for Western Armenian, an endangered and low-resource Indo-European language. The transducer has virtually complete coverage of the language’s inflectional morphology…

Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press

2026-08-06 · Chahan Vidal-Gorène, Seda Kirakosyan, Edita Matevosyan arxiv

This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Parisian Armenian commercial community. On each page, commercial advert…

AthDGC: An Open Diachronic Greek Treebank with Indo-European Parallels

2026-06-13 · Nikolaos Lavidas, Kiki Nikiforidou, Dag Haug, Leonid Kulikov 외 arxiv

AthDGC ("Athens-PROIEL") is an open, end-to-end workflow and dataset. It is, to the best of our knowledge, the first openly licensed dependency-parsed treebank of Greek that spans eight diachronic periods, namely Archaic…

Dialects Identification of Armenian Language

2022-06-01 · DigitAm (LREC) 2022 6 · Karen Avetisyan

The Armenian language has many dialects that differ from each other syntactically, morphologically, and phonetically. In this work, we implement and evaluate models that determine the dialect of a given passage of text. …

Dialect IdentificationLanguage IdentificationSentenceWord Embeddings