paper-with-me

Papers

Low-Resource Machine Transliteration Using Recurrent Neural Networks of Asian Languages

2018-07-01 · WS 2018 7 · Ngoc Tan Le, Fatiha Sadat

Grapheme-to-phoneme models are key components in automatic speech recognition and text-to-speech systems. With low-resource language pairs that do not have available and well-developed pronunciation lexicons, grapheme-to-phoneme models are particularly useful. These models are based on initial alignments between grapheme source and phoneme target sequences. Inspired by sequence-to-sequence recurrent neural network-based translation methods, the current research presents an approach that applies an alignment representation for input sequences and pre-trained source and target embeddings to overcome the transliteration problem for a low-resource languages pair. We participated in the NEWS 2018 shared task for the English-Vietnamese transliteration task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognitionSpeech Recognitiontext-to-speechText to SpeechTranslationTransliterationWord Embeddings

Similar Papers 제목 키워드 기반

Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset

2020-07-02 · LREC 2020 5 · Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke 외

This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each language: 1) native script Wikipedia text; 2)…

Language ModelingLanguage ModellingSentenceTransliteration

Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment

2024-06-28 · Orgest Xhelili, Yihong Liu, Hinrich Schütze

Multilingual pre-trained models (mPLMs) have shown impressive performance on cross-lingual transfer tasks. However, the transfer performance is often hindered when a low-resource target language is written in a different…

Cross-Lingual TransferTransliteration

Speech Synthesis for Low Resource Languages using Transliteration Enabled Transfer Learning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In the area of Human Computer Interaction (HCI), Text To Speech (TTS) synthesis has received a significant boost in recent years, especially with the development of various deep learning techniques capable of generating …

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+3

Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration

2025-11-27 · Kanchon Gharami, Quazi Sarwar Muhtaseem, Deepti Gupta, Lavanya Elluri 외 arxiv

The development of robust transliteration techniques to enhance the effectiveness of transforming Romanized scripts into native scripts is crucial for Natural Language Processing tasks, including sentiment analysis, spee…

Information RetrievalSpeech RecognitionSentiment Analysis

Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages

2025-09-22 · Wenhao Zhuang, Yuan Sun, Xiaobing Zhao arxiv

As large language models (LLMs) are trained on increasingly diverse and extensive multilingual corpora, they demonstrate cross-lingual transfer capabilities. However, these capabilities often fail to effectively extend t…

Machine Reading ComprehensionCross-Lingual TransferMachine TranslationText Classification