Bridging the Language Gap: Synthetic Voice Diversity via Latent Mixup for Equitable Speech Recognition
Modern machine learning models for audio tasks often exhibit superior performance on English and other well-resourced languages, primarily due to the abundance of available training data. This disparity leads to an unfair performance gap for low-resource languages, where data collection is both challenging and costly. In this work, we introduce a novel data augmentation technique for speech corpora designed to mitigate this gap. Through comprehensive experiments, we demonstrate that our method significantly improves the performance of automatic speech recognition systems on low-resource languages. Furthermore, we show that our approach outperforms existing augmentation strategies, offering a practical solution for enhancing speech technology in underrepresented linguistic communities.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech RecognitionData AugmentationSimilar Papers 제목 키워드 기반
Using Audio Books for Training a Text-to-Speech System
Creating new voices for a TTS system often requires a costly procedure of designing and recording an audio corpus, a time consuming and effort intensive task. Using publicly available audiobooks as the raw material of a …
DiversitySpeech Synthesistext-to-speechText to SpeechPromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts
Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or ref…
Voice ConversionVoiceAgentBench: Are Voice Assistants ready for agentic tasks?
Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities su…
Adversarial RobustnessQuestion AnsweringVoice ConversionVoice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script wi…
DiversityVoice ConversionVOLTA: Improving Generative Diversity by Variational Mutual Information Maximizing Autoencoder
The natural language generation domain has witnessed great success thanks to Transformer models. Although they have achieved state-of-the-art generative quality, they often neglect generative diversity. Prior attempts to…
DiversityText Generation