paper-with-me

홈 › Papers

Morphological Analysis Corpus Construction of Uyghur

2021-08-01 · CCL 2021 8 · Abudouwaili Gulinigeer, Abiderexiti Kahaerjiang, Wushouer Jiamila, Shen Yunfei, Maimaitimin Turenisha, Yibulayin Tuergen

“Morphological analysis is a fundamental task in natural language processing and results can beapplied to different downstream tasks such as named entity recognition syntactic analysis andmachine translation. However there are many problems in morphological analysis such as lowaccuracy caused by a lack of resources. In this paper to alleviate the lack of resources in Uyghurmorphological analysis research we construct a Uyghur morphological analysis corpus based onthe analysis of grammatical features and the format of the general morphological analysis corpus.We define morphological tags from 14 dimensions and 53 features manually annotate and correctthe dataset. Finally the corpus provided some informations such as word lemma part of speech morphological analysis tags morphological segmentation and lemmatization. Also this paperanalyzes some basic features of the corpus and we use the models and datasets provided bySIGMORPHON Shared Task organizers to design comparative experiments to verify the corpus’savailability. Results of the experiment are 85.56% 88.29% respectively. The corpus provides areference value for morphological analysis and promotes the research of Uyghur natural language processing.”

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

LEMMALemmatizationMorphological Analysisnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Translation

Similar Papers 제목 키워드 기반

Log-linear Models for Uyghur Segmentation in Spoken Language Translation

2017-09-01 · RANLP 2017 9 · Chenggang Mi, Yating Yang, Rui Dong, Xi Zhou 외

To alleviate data sparsity in spoken Uyghur machine translation, we proposed a log-linear based morphological segmentation approach. Instead of learning model only from monolingual annotated corpus, this approach optimiz…

Machine TranslationSegmentationTranslationWord Alignment

Building Language Models for Morphological Rich Low-Resource Languages using Data from Related Donor Languages: the Case of Uyghur

2020-05-01 · LREC 2020 5 · Ayimunishagu Abulimiti, Tanja Schultz

Huge amounts of data are needed to build reliable statistical language models. Automatic speech processing tasks in low-resource languages typically suffer from lower performances due to weak or unreliable language model…

Language ModelingLanguage Modelling

UgChDial: A Uyghur Chat-based Dialogue Corpus for Response Space Classification

2022-06-01 · LREC 2022 6 · Zulipiye Yusupujiang, Jonathan Ginzburg

In this paper, we introduce a carefully designed and collected language resource: UgChDial – a Uyghur dialogue corpus based on a chatroom environment. The Uyghur Chat-based Dialogue Corpus (UgChDial) is divided into two …

Constructing Uyghur Name Entity Recognition System using Neural Machine Translation Tag Projection

2020-10-01 · CCL 2020 10 · Anwar Azmat, Li Xiao, Yang Yating, Dong Rui 외

Although named entity recognition achieved great success by introducing the neural networks, it is challenging to apply these models to low resource languages including Uyghur while it depends on a large amount of annota…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4

CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages

2025-09-21 · Wenhao Zhuang, Yuan Sun arxiv

Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich lang…

Cross-Lingual TransferMachine Translation