Introducing YakuToolkit. Yakut Treebank and Morphological Analyzer.
This poster presents the first publicly available treebank of Yakut, a Turkic language spoken in Russia, and a morphological analyzer for this language. The treebank was annotated following the Universal Dependencies (UD) framework and the mor- phological analyzer can directly access and use its data. Yakut is an under-represented language whose prominence can be raised by making reliably annotated data and NLP tools that could process it freely accessible. The publication of both the treebank and the analyzer serves this purpose with the prospect of evolving into a benchmark for the development of NLP online tools for other languages of the Turkic family in the future.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Preparing Korean Data for the Shared Task on Parsing Morphologically Rich Languages
This document gives a brief description of Korean data prepared for the SPMRL 2013 shared task. A total of 27,363 sentences with 350,090 tokens are used for the shared task. All constituent trees are collected from the K…
Morphological AnalysisDeveloping an Egyptian Arabic Treebank: Impact of Dialectal Morphology on Annotation and Tool Development
This paper describes the parallel development of an Egyptian Arabic Treebank and a morphological analyzer for Egyptian Arabic (CALIMA). By the very nature of Egyptian Arabic, the data collected is informal, for example D…
The Impact of Automatic Morphological Analysis \& Disambiguation on Dependency Parsing of Turkish
The studies on dependency parsing of Turkish so far gave their results on the Turkish Dependency Treebank. This treebank consists of sentences where gold standard part-of-speech tags are manually assigned to each word an…
Dependency ParsingInformation RetrievalMorphological AnalysisMorphological Disambiguation of South Sámi with FSTs and Neural Networks
We present a method for conducting morphological disambiguation for South S\'ami, which is an endangered language. Our method uses an FST-based morphological analyzer to produce an ambiguous set of morphological readings…
Morphological DisambiguationSentenceWord EmbeddingsMorphological Disambiguation of South S\'ami with FSTs and Neural Networks
We present a method for conducting morphological disambiguation for South S{\'a}mi, which is an endangered language. Our method uses an FST-based morphological analyzer to produce an ambiguous set of morphological readin…
Morphological DisambiguationSentenceWord Embeddings