paper-with-me

홈 › Papers

Noisy Uyghur Text Normalization

2017-09-01 · WS 2017 9 · Osman Tursun, Ruket Cakici

Uyghur is the second largest and most actively used social media language in China. However, a non-negligible part of Uyghur text appearing in social media is unsystematically written with the Latin alphabet, and it continues to increase in size. Uyghur text in this format is incomprehensible and ambiguous even to native Uyghur speakers. In addition, Uyghur texts in this form lack the potential for any kind of advancement for the NLP tasks related to the Uyghur language. Restoring and preventing noisy Uyghur text written with unsystematic Latin alphabets will be essential to the protection of Uyghur language and improving the accuracy of Uyghur NLP tasks. To this purpose, in this work we propose and compare the noisy channel model and the neural encoder-decoder model as normalizing methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderText Normalization

Similar Papers 제목 키워드 기반

Automatic Speech Recognition for Uyghur through Multilingual Acoustic Modeling

2020-05-01 · LREC 2020 5 · Ayimunishagu Abulimiti, Tanja Schultz

Low-resource languages suffer from lower performance of Automatic Speech Recognition (ASR) system due to the lack of data. As a common approach, multilingual training has been applied to achieve more context coverage and…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Building Language Models for Morphological Rich Low-Resource Languages using Data from Related Donor Languages: the Case of Uyghur

2020-05-01 · LREC 2020 5 · Ayimunishagu Abulimiti, Tanja Schultz

Huge amounts of data are needed to build reliable statistical language models. Automatic speech processing tasks in low-resource languages typically suffer from lower performances due to weak or unreliable language model…

Language ModelingLanguage Modelling

Toward Better Loanword Identification in Uyghur Using Cross-lingual Word Embeddings

2018-08-01 · COLING 2018 8 · Chenggang Mi, Yating Yang, Lei Wang, Xi Zhou 외

To enrich vocabulary of low resource settings, we proposed a novel method which identify loanwords in monolingual corpora. More specifically, we first use cross-lingual word embeddings as the core feature to generate sem…

Cross-Lingual Word EmbeddingsLanguage ModelingLanguage ModellingMachine Translation+2

Universal dependencies for Uyghur

2016-12-01 · WS 2016 12 · Marhaba Eli, Weinila Mushajiang, Tuergen Yibulayin, Kahaerjiang Abiderexiti 외

The Universal Dependencies (UD) Project seeks to build a cross-lingual studies of treebanks, linguistic structures and parsing. Its goal is to create a set of multilingual harmonized treebanks that are designed according…

Cross-Lingual Transfer

Memory-augmented Chinese-Uyghur Neural Machine Translation

2017-06-27 · Shiyue Zhang, Gulnigar Mahmut, Dong Wang, Askar Hamdulla

Neural machine translation (NMT) has achieved notable performance recently. However, this approach has not been widely applied to the translation task between Chinese and Uyghur, partly due to the limited parallel data r…

Machine TranslationNMTSentenceTranslation