paper-with-me

홈 › Papers

Grapheme-to-Phoneme Transformer Model for Transfer Learning Dialects

2021-04-08 · Eric Engelhart, Mahsa Elyasi, Gaurav Bharaj

Grapheme-to-Phoneme (G2P) models convert words to their phonetic pronunciations. Classic G2P methods include rule-based systems and pronunciation dictionaries, while modern G2P systems incorporate learning, such as, LSTM and Transformer-based attention models. Usually, dictionary-based methods require significant manual effort to build, and have limited adaptivity on unseen words. And transformer-based models require significant training data, and do not generalize well, especially for dialects with limited data. We propose a novel use of transformer-based attention model that can adapt to unseen dialects of English language, while using a small dictionary. We show that our method has potential applications for accent transfer for text-to-speech, and for building robust G2P models for dialects with limited pronunciation dictionary size. We experiment with two English dialects: Indian and British. A model trained from scratch using 1000 words from British English dictionary, with 14211 words held out, leads to phoneme error rate (PER) of 26.877%, on a test set generated using the full dictionary. The same model pretrained on CMUDict American English dictionary, and fine-tuned on the same dataset leads to PER of 2.469% on the test set.

📄 PDF Abstract BibTeX arXiv:2104.04091

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to SpeechTransfer Learning

Methods 이 논문이 사용한 방법론

American 설명 없음
Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

A Swiss German Dictionary: Variation in Speech and Writing

2020-03-31 · LREC 2020 5 · Larissa Schmidt, Lucy Linder, Sandra Djambazovska, Alexandros Lazaridis 외

We introduce a dictionary containing forms of common words in various Swiss German dialects normalized into High German. As Swiss German is, for now, a predominantly spoken language, there is a significant variation in t…

Diversityspeech-recognitionSpeech RecognitionTranslation

No Need for a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End Models

2017-12-05 · Tara N. Sainath, Rohit Prabhavalkar, Shankar Kumar, Seungji Lee 외

For decades, context-dependent phonemes have been the dominant sub-word unit for conventional acoustic modeling systems. This status quo has begun to be challenged recently by end-to-end models which seek to combine acou…

Language ModelingLanguage Modelling

Optimizing Two-Pass Cross-Lingual Transfer Learning: Phoneme Recognition and Phoneme to Grapheme Translation

2023-12-06 · Wonjun Lee, Gary Geunbae Lee, Yunsu Kim

This research optimizes two-pass cross-lingual transfer learning in low-resource languages by enhancing phoneme recognition and phoneme-to-grapheme translation models. Our approach optimizes these two stages to improve s…

Cross-Lingual TransferPhoneme Recognitionspeech-recognitionSpeech Recognition+1

DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation

2025-09-25 · Ziqi Chen, Gongyu Chen, Yihua Wang, Chaofan Ding 외 arxiv

Dialect speech embodies rich cultural and linguistic diversity, yet building text-to-speech (TTS) systems for dialects remains challenging due to scarce data, inconsistent orthographies, and complex phonetic variation. T…

Multilingual Speech Recognition for Low-Resource Indian Languages using Multi-Task conformer

2021-08-22 · Krishna D N

Transformers have recently become very popular for sequence-to-sequence applications such as machine translation and speech recognition. In this work, we propose a multi-task learning-based transformer model for low-reso…

DecoderMachine TranslationMulti-Task LearningPhoneme Recognition+3