paper-with-me

Papers

Text is All You Need: Personalizing ASR Models using Controllable Speech Synthesis

2023-03-27 · Karren Yang, Ting-yao Hu, Jen-Hao Rick Chang, Hema Swetha Koppula, Oncel Tuzel

Adapting generic speech recognition models to specific individuals is a challenging problem due to the scarcity of personalized data. Recent works have proposed boosting the amount of training data using personalized text-to-speech synthesis. Here, we ask two fundamental questions about this strategy: when is synthetic data effective for personalization, and why is it effective in those cases? To address the first question, we adapt a state-of-the-art automatic speech recognition (ASR) model to target speakers from four benchmark datasets representative of different speaker types. We show that ASR personalization with synthetic data is effective in all cases, but particularly when (i) the target speaker is underrepresented in the global data, and (ii) the capacity of the global model is limited. To address the second question of why personalized synthetic data is effective, we use controllable speech synthesis to generate speech with varied styles and content. Surprisingly, we find that the text content of the synthetic data, rather than style, is important for speaker adaptation. These results lead us to propose a data selection strategy for ASR personalization based on speech content.

📄 PDF Abstract BibTeX arXiv:2303.14885

Code (0)

등록된 구현이 없습니다.

Tasks

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis

2025-05-25 · Minsu Kim, Pingchuan Ma, Honglie Chen, Stavros Petridis 외

This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

2024-12-09 · Tianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang 외

Text-to-speech (TTS), also known as speech synthesis, is a prominent research area that aims to generate natural-sounding human speech from text. Recently, with the increasing industrial demand, TTS technologies have evo…

Speech SynthesisSurveytext-to-speechText to Speech

Corpus Synthesis for Zero-shot ASR domain Adaptation using Large Language Models

2023-09-18 · Hsuan Su, Ting-yao Hu, Hema Swetha Koppula, Raviteja Vemulapalli 외

While Automatic Speech Recognition (ASR) systems are widely used in many real-world applications, they often do not generalize well to new domains and need to be finetuned on data from these domains. However, target-doma…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain AdaptationLanguage Modeling+5

Learning Robust Latent Representations for Controllable Speech Synthesis

2021-05-10 · Findings (ACL) 2021 8 · Shakti Kumar, Jithin Pradeep, Hussain Zaidi

State-of-the-art Variational Auto-Encoders (VAEs) for learning disentangled latent representations give impressive results in discovering features like pitch, pause duration, and accent in speech data, leading to highly …

Speech Synthesistext-to-speechText to Speech

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

2025-01-08 · Hanzhao Li, Yuke Li, Xinsheng Wang, Jingbin Hu 외

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjust…

Speech Synthesis