paper-with-me

Papers

G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR

2019-10-22 · Duc Le, Thilo Koehler, Christian Fuegen, Michael L. Seltzer

Grapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonemic languages like English. However, graphemic ASR still has problems with rare long-tail words that do not follow the standard spelling conventions seen in training, such as entity names. In this work, we present a novel method to train a statistical grapheme-to-grapheme (G2G) model on text-to-speech data that can rewrite an arbitrary character sequence into more phonetically consistent forms. We show that using G2G to provide alternative pronunciations during decoding reduces Word Error Rate by 3% to 11% relative over a strong graphemic baseline and bridges the gap on rare name recognition with an equivalent phonetic setup. Unlike many previously proposed methods, our method does not require any change to the acoustic model training procedure. This work reaffirms the efficacy of grapheme-based modeling and shows that specialized linguistic knowledge, when available, can be leveraged to improve graphemic ASR.

📄 PDF Abstract BibTeX arXiv:1910.12612

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Neural TTS in French: Comparing Graphemic and Phonetic Inputs Using the SynPaFlex-Corpus and Tacotron2

2023-04-17 · Samuel Delalez, Ludi Akue

The SynPaFlex-Corpus is a publicly available TTS-oriented dataset, which provides phonetic transcriptions automatically produced by the JTrans transcriber, with a Phoneme Error Rate (PER) of 6.1%. In this paper, we analy…

Multimodal neural pronunciation modeling for spoken languages with logographic origin

2018-09-12 · EMNLP 2018 10 · Minh Nguyen, Gia H. Ngo, Nancy F. Chen

Graphemes of most languages encode pronunciation, though some are more explicit than others. Languages like Spanish have a straightforward mapping between its graphemes and phonemes, while this mapping is more convoluted…

Deep Shallow Fusion for RNN-T Personalization

2020-11-16 · Duc Le, Gil Keren, Julian Chan, Jay Mahadeokar 외

End-to-end models in general, and Recurrent Neural Network Transducer (RNN-T) in particular, have gained significant traction in the automatic speech recognition community in the last few years due to their simplicity, c…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Multilingual Graphemic Hybrid ASR with Massive Data Augmentation

2019-09-14 · LREC 2020 5 · Chunxi Liu, Qiaochu Zhang, Xiaohui Zhang, Kritika Singh 외

Towards developing high-performing ASR for low-resource languages, approaches to address the lack of resources are to make use of data from multiple languages, and to augment the training data by creating acoustic variat…

Data Augmentation

Comparison of Grapheme-to-Phoneme Conversion Methods on a Myanmar Pronunciation Dictionary

2016-12-01 · WS 2016 12 · Ye Kyaw Thu, Win Pa Pa, Yoshinori Sagisaka, Naoto Iwahashi

Grapheme-to-Phoneme (G2P) conversion is the task of predicting the pronunciation of a word given its graphemic or written form. It is a highly important part of both automatic speech recognition (ASR) and text-to-speech …

Active LearningAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Grapheme-to-Phoneme Conversion+6