Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
Currently, a common approach in many speech processing tasks is to leverage large scale pre-trained models by fine-tuning them on in-domain data for a particular application. Yet obtaining even a small amount of such data can be problematic, especially for sensitive domains and conversational speech scenarios, due to both privacy issues and annotation costs. To address this, synthetic data generation using single speaker datasets has been employed. Yet, for multi-speaker cases, such an approach often requires extensive manual effort and is prone to domain mismatches. In this work, we propose a synthetic data generation pipeline for multi-speaker conversational ASR, leveraging a large language model (LLM) for content creation and a conversational multi-speaker text-to-speech (TTS) model for speech synthesis. We conduct evaluation by fine-tuning the Whisper ASR model for telephone and distant conversational speech settings, using both in-domain data and generated synthetic data. Our results show that the proposed method is able to significantly outperform classical multi-speaker generation approaches that use external, non-conversational speech datasets.
Code (1)
Tasks
Language ModelingLanguage ModellingLarge Language Modelspeech-recognitionSpeech RecognitionSpeech SynthesisSynthetic Data Generationtext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
Recently, there has been a growing interest in the field of controllable Text-to-Speech (TTS). While previous studies have relied on users providing specific style factor values based on acoustic knowledge or selecting r…
Language Modellingtext-to-speechText to SpeechJELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
Recently, there has been a growing demand for conversational speech synthesis (CSS) that generates more natural speech by considering the conversational context. To address this, we introduce JELLY, a novel CSS framework…
Emotion RecognitionLanguage ModelingLanguage ModellingLarge Language Model+1Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. Addressing the scarcity of samples contai…
Synthetic Data Generationtext-to-speechText to SpeechOutcome-Constrained Large Language Models for Countering Hate Speech
Automatic counterspeech generation methods have been developed to assist efforts in combating hate speech. Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informativ…
Reinforcement Learning (RL)Text GenerationVITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
With the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+6