paper-with-me

Papers

Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech

2024-01-19 · Abhinav Garg, Jiyeon Kim, Sushil Khyalia, Chanwoo Kim, Dhananjaya Gowda

Grapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, usually, ARPABET or IPA, which might not be the most optimal way to represent phonemes for all languages. Secondly, the man-hours required to produce such an expert lexicon are very high. In this paper, we eliminate both of these issues by using recent advances in self-supervised learning to obtain data-driven phoneme representations instead of fixed representations. We compare our lexicon-free approach against strong baselines that utilize a well-crafted lexicon. Furthermore, we show that our data-driven lexicon-free method performs as good or even marginally better than the conventional rule-based or lexicon-based neural G2Ps in terms of Mean Opinion Score (MOS) while using no prior language lexicon or phoneme set, i.e. no linguistic expertise.

📄 PDF Abstract BibTeX arXiv:2401.10465

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learningtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Customizing Grapheme-to-Phoneme System for Non-Trivial Transcription Problems in Bangla Language

2019-06-01 · NAACL 2019 6 · Sudipta Saha Shubha, Nafis Sadeq, Shafayat Ahmed, Md. Nahidul Islam 외

Grapheme to phoneme (G2P) conversion is an integral part in various text and speech processing systems, such as: Text to Speech system, Speech Recognition system, etc. The existing methodologies for G2P conversion in Ban…

speech-recognitionSpeech Recognitiontext-to-speechText to Speech

No Need for a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End Models

2017-12-05 · Tara N. Sainath, Rohit Prabhavalkar, Shankar Kumar, Seungji Lee 외

For decades, context-dependent phonemes have been the dominant sub-word unit for conventional acoustic modeling systems. This status quo has begun to be challenged recently by end-to-end models which seek to combine acou…

Language ModelingLanguage Modelling

LSTM Acoustic Models Learn to Align and Pronounce with Graphemes

2020-08-13 · Arindrima Datta, Guanlong Zhao, Bhuvana Ramabhadran, Eugene Weinstein

Automated speech recognition coverage of the world's languages continues to expand. However, standard phoneme based systems require handcrafted lexicons that are difficult and expensive to obtain. To address this problem…

speech-recognitionSpeech Recognition

Mlphon: A Multifunctional Grapheme-Phoneme Conversion Tool Using Finite State Transducers

2022-09-05 · IEEE Access 2022 9 · Kavya Manohar, A R jayan, Rajeev Rajan

In this article we present the design and the development of a knowledge based computational linguistic tool, Mlphon for Malayalam language. Mlphon computationally models linguistic rules using finite state transducers a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityGrapheme-to-Phoneme Conversion+10

Low-Resource Machine Transliteration Using Recurrent Neural Networks of Asian Languages

2018-07-01 · WS 2018 7 · Ngoc Tan Le, Fatiha Sadat

Grapheme-to-phoneme models are key components in automatic speech recognition and text-to-speech systems. With low-resource language pairs that do not have available and well-developed pronunciation lexicons, grapheme-to…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+6