paper-with-me

Papers

A machine transliteration tool between Uzbek alphabets

2022-05-19 · Ulugbek Salaev, Elmurod Kuriyozov, Carlos Gómez-Rodríguez

Machine transliteration, as defined in this paper, is a process of automatically transforming written script of words from a source alphabet into words of another target alphabet within the same language, while preserving their meaning, as well as pronunciation. The main goal of this paper is to present a machine transliteration tool between three common scripts used in low-resource Uzbek language: the old Cyrillic, currently official Latin, and newly announced New Latin alphabets. The tool has been created using a combination of rule-based and fine-tuning approaches. The created tool is available as an open-source Python package, as well as a web-based application including a public API. To our knowledge, this is the first machine transliteration tool that supports the newly announced Latin alphabet of the Uzbek language.

📄 PDF Abstract BibTeX arXiv:2205.09578

Code (1)

ulugbeksalaev/uztransliterator 공식 구현

Tasks

Transliteration

Similar Papers 제목 키워드 기반

Uzbek Cyrillic-Latin-Cyrillic Machine Transliteration

2021-01-13 · B. Mansurov, A. Mansurov

In this paper, we introduce a data-driven approach to transliterating Uzbek dictionary words from the Cyrillic script into the Latin script, and vice versa. We heuristically align characters of words in the source script…

Transliteration

FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus with Punctuation, Casing, and Context

2024-09-23 · Anna Povey, Katherine Povey

This paper introduces FeruzaSpeech, a read speech corpus of the Uzbek language, containing transcripts in both Cyrillic and Latin alphabets, freely available for academic research purposes. This corpus includes 60 hours …

Multilingual Text-to-Speech Synthesis for Turkic Languages Using Transliteration

2023-05-25 · Rustem Yeshpanov, Saida Mussakhojayeva, Yerbolat Khassanov

This work aims to build a multilingual text-to-speech (TTS) synthesis system for ten lower-resourced Turkic languages: Azerbaijani, Bashkir, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Turkmen, Uyghur, and Uzbek. We specifica…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+2

Context based Roman-Urdu to Urdu Script Transliteration System

2021-09-29 · H Muhammad Shakeel, Rashid Khan, Muhammad Waheed

Now a day computer is necessary for human being and it is very useful in many fields like search engine, text processing, short messaging services, voice chatting and text recognition. Since last many years there are man…

Transliteration

A complete character recognition and transliteration technique for Devanagari script

2020-09-28 · Jasmine Kaur, Vinay Kumar

Transliteration involves transformation of one script to another based on phonetic similarities between the characters of two distinctive scripts. In this paper, we present a novel technique for automatic transliteration…

SegmentationTransliteration