Fast Bootstrapping of Grapheme to Phoneme System for Under-resourced Languages - Application to the Iban Language
Code (0)
등록된 구현이 없습니다.
Tasks
Speech RecognitionSpeech SynthesisText-To-Speech SynthesisSimilar Papers 제목 키워드 기반
Comparison of Grapheme-to-Phoneme Conversion Methods on a Myanmar Pronunciation Dictionary
Grapheme-to-Phoneme (G2P) conversion is the task of predicting the pronunciation of a word given its graphemic or written form. It is a highly important part of both automatic speech recognition (ASR) and text-to-speech …
Active LearningAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Grapheme-to-Phoneme Conversion+6Fast Bilingual Grapheme-To-Phoneme Conversion
Autoregressive transformer (ART)-based grapheme-to-phoneme (G2P) models have been proposed for bi/multilingual text-to-speech systems. Although they have achieved great success, they suffer from high inference latency in…
Data AugmentationGrapheme-to-Phoneme ConversionSentencetext-to-speech+1No Need for a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End Models
For decades, context-dependent phonemes have been the dominant sub-word unit for conventional acoustic modeling systems. This status quo has begun to be challenged recently by end-to-end models which seek to combine acou…
Language ModelingLanguage ModellingA systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models
Following the rationale of end-to-end modeling, CTC, RNN-T or encoder-decoder-attention models for automatic speech recognition (ASR) use graphemes or grapheme-based subword units based on e.g. byte-pair encoding (BPE). …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1An Investigation of the Relation Between Grapheme Embeddings and Pronunciation for Tacotron-based Systems
End-to-end models, particularly Tacotron-based ones, are currently a popular solution for text-to-speech synthesis. They allow the production of high-quality synthesized speech with little to no text preprocessing. Indee…
Grapheme-to-Phoneme ConversionRelationSpeech Synthesistext-to-speech+2