paper-with-me

홈 › Papers

Benchmarking Azerbaijani Neural Machine Translation

2022-07-29 · Chih-Chen Chen, William Chen

Little research has been done on Neural Machine Translation (NMT) for Azerbaijani. In this paper, we benchmark the performance of Azerbaijani-English NMT systems on a range of techniques and datasets. We evaluate which segmentation techniques work best on Azerbaijani translation and benchmark the performance of Azerbaijani NMT models across several domains of text. Our results show that while Unigram segmentation improves NMT performance and Azerbaijani translation models scale better with dataset quality than quantity, cross-domain generalization remains a challenge

📄 PDF Abstract BibTeX arXiv:2207.14473

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDomain GeneralizationMachine TranslationNMTSegmentationTranslation

Methods 이 논문이 사용한 방법론

Unigram Segmentation Unigram Segmentation is a subword segmentation algorithm based on a unigram language model. It provides multiple segmentations with probabilities. The language model allows…

Similar Papers 제목 키워드 기반

Neural machine translation system for Lezgian, Russian and Azerbaijani languages

2024-10-07 · Alidar Asvarov, Andrey Grabovoy

We release the first neural machine translation system for translation between Russian, Azerbaijani and the endangered Lezgian languages, as well as monolingual and parallel datasets collected and aligned for training an…

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+2

Open foundation models for Azerbaijani language

2024-07-02 · Jafar Isbarov, Kavsar Huseynova, Elvin Mammadov, Mammad Hajili 외

The emergence of multilingual large language models has enabled the development of language understanding and generation systems in Azerbaijani. However, most of the production-grade systems rely on cloud solutions, such…

Benchmarking

Enhancing Language Learning through Technology: Introducing a New English-Azerbaijani (Arabic Script) Parallel Corpus

2024-07-06 · Jalil Nourmohammadi Khiarak, Ammar Ahmadi, Taher Ak-bari Saeed, Meysam Asgari-Chenaghlu 외

This paper introduces a pioneering English-Azerbaijani (Arabic Script) parallel corpus, designed to bridge the technological gap in language learning and machine translation (MT) for under-resourced languages. Consisting…

ArticlesMachine TranslationNMTTranslation

Text Classification for Azerbaijani Language Using Machine Learning and Embedding

2019-12-26 · Umid Suleymanov, Behnam Kiani Kalejahi, Elkhan Amrahov, Rashid Badirkhanli

Text classification systems will help to solve the text clustering problem in the Azerbaijani language. There are some text-classification applications for foreign languages, but we tried to build a newly developed syste…

BIG-bench Machine LearningClassificationClusteringGeneral Classification+4

AzSLD: Azerbaijani Sign Language Dataset for Fingerspelling, Word, and Sentence Translation with Baseline Software

2024-11-19 · Nigar Alishzade, Jamaladdin Hasanov

Sign language processing technology development relies on extensive and reliable datasets, instructions, and ethical guidelines. We present a comprehensive Azerbaijani Sign Language Dataset (AzSLD) collected from diverse…

Gesture RecognitionSentenceSign Language RecognitionTranslation