paper-with-me

홈 › Papers

Log-linear Models for Uyghur Segmentation in Spoken Language Translation

2017-09-01 · RANLP 2017 9 · Chenggang Mi, Yating Yang, Rui Dong, Xi Zhou, Lei Wang, Xiao Li, Tonghai Jiang

To alleviate data sparsity in spoken Uyghur machine translation, we proposed a log-linear based morphological segmentation approach. Instead of learning model only from monolingual annotated corpus, this approach optimizes Uyghur segmentation for spoken translation based on both bilingual and monolingual corpus. Our approach relies on several features such as traditional conditional random field (CRF) feature, bilingual word alignment feature and monolingual suffixword co-occurrence feature. Experimental results shown that our proposed segmentation model for Uyghur spoken translation achieved 1.6 BLEU score improvements compared with the state-of-the-art baseline.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSegmentationTranslationWord Alignment

Similar Papers 제목 키워드 기반

Toward Better Loanword Identification in Uyghur Using Cross-lingual Word Embeddings

2018-08-01 · COLING 2018 8 · Chenggang Mi, Yating Yang, Lei Wang, Xi Zhou 외

To enrich vocabulary of low resource settings, we proposed a novel method which identify loanwords in monolingual corpora. More specifically, we first use cross-lingual word embeddings as the core feature to generate sem…

Cross-Lingual Word EmbeddingsLanguage ModelingLanguage ModellingMachine Translation+2

Morphological Analysis Corpus Construction of Uyghur

2021-08-01 · CCL 2021 8 · Abudouwaili Gulinigeer, Abiderexiti Kahaerjiang, Wushouer Jiamila, Shen Yunfei 외

“Morphological analysis is a fundamental task in natural language processing and results can beapplied to different downstream tasks such as named entity recognition syntactic analysis andmachine translation. However the…

LEMMALemmatizationMorphological Analysisnamed-entity-recognition+3

Memory-augmented Chinese-Uyghur Neural Machine Translation

2017-06-27 · Shiyue Zhang, Gulnigar Mahmut, Dong Wang, Askar Hamdulla

Neural machine translation (NMT) has achieved notable performance recently. However, this approach has not been widely applied to the translation task between Chinese and Uyghur, partly due to the limited parallel data r…

Machine TranslationNMTSentenceTranslation

CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages

2025-09-21 · Wenhao Zhuang, Yuan Sun arxiv

Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich lang…

Cross-Lingual TransferMachine Translation

Constructing Uyghur Name Entity Recognition System using Neural Machine Translation Tag Projection

2020-10-01 · CCL 2020 10 · Anwar Azmat, Li Xiao, Yang Yating, Dong Rui 외

Although named entity recognition achieved great success by introducing the neural networks, it is challenging to apply these models to low resource languages including Uyghur while it depends on a large amount of annota…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4