paper-with-me

Papers

Creating a morphological and syntactic tagged corpus for the Uzbek language

2022-10-27 · Maksud Sharipov, Jamolbek Mattiev, Jasur Sobirov, Rustam Baltayev

Nowadays, creation of the tagged corpora is becoming one of the most important tasks of Natural Language Processing (NLP). There are not enough tagged corpora to build machine learning models for the low-resource Uzbek language. In this paper, we tried to fill that gap by developing a novel Part Of Speech (POS) and syntactic tagset for creating the syntactic and morphologically tagged corpus of the Uzbek language. This work also includes detailed description and presentation of a web-based application to work on a tagging as well. Based on the developed annotation tool and the software, we share our experience results of the first stage of the tagged corpus creation

📄 PDF Abstract BibTeX arXiv:2210.15234

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Similar Papers 제목 키워드 기반

Contemporary Amharic Corpus: Automatically Morpho-Syntactically Tagged Amharic Corpus

2021-06-14 · COLING 2018 8 · Andargachew Mekonnen Gezmu, Binyam Ephrem Seyoum, Michael Gasser, Andreas Nürnberger

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are …

Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek

2025-08-20 · Mukhammadsaid Mamasaidov, Azizullah Aral, Abror Shopulatov, Mironshoh Inomjonov arxiv

Southern Uzbek (uzs) is a Turkic language variety spoken by around 5 million people in Afghanistan and differs significantly from Northern Uzbek (uzn) in phonology, lexicon, and orthography. Despite the large number of s…

Machine Translation

BBPOS: BERT-based Part-of-Speech Tagging for Uzbek

2025-01-17 · Latofat Bobojonova, Arofat Akhundjanova, Phil Ostheimer, Sophie Fellenz

This paper advances NLP research for the low-resource Uzbek language by evaluating two previously untested monolingual Uzbek BERT models on the part-of-speech (POS) tagging task and introducing the first publicly availab…

Part-Of-Speech TaggingPOSPOS TaggingSensitivity

Accuracy of the Uzbek stop words detection: a case study on "School corpus"

2022-09-15 · Khabibulla Madatov, Shukurla Bekchanov, Jernej Vičič

Stop words are very important for information retrieval and text analysis investigation tasks of natural language processing. Current work presents a method to evaluate the quality of a list of stop words aimed at automa…

Information RetrievalRetrievalSentence

Szeged Corpus 2.5: Morphological Modifications in a Manually POS-tagged Hungarian Corpus

2014-05-01 · LREC 2014 5 · Veronika Vincze, Viktor Varga, Katalin Ilona Simk{\'o}, J{\'a}nos Zsibrita 외

The Szeged Corpus is the largest manually annotated database containing the possible morphological analyses and lemmas for each word form. In this work, we present its latest version, Szeged Corpus 2.5, in which the new …

POS