paper-with-me

홈 › Papers

Bridging the Language Gap: Synthetic Voice Diversity via Latent Mixup for Equitable Speech Recognition

2025-11-25 · Wesley Bian, Xiaofeng Lin, Guang Cheng arxiv

Modern machine learning models for audio tasks often exhibit superior performance on English and other well-resourced languages, primarily due to the abundance of available training data. This disparity leads to an unfair performance gap for low-resource languages, where data collection is both challenging and costly. In this work, we introduce a novel data augmentation technique for speech corpora designed to mitigate this gap. Through comprehensive experiments, we demonstrate that our method significantly improves the performance of automatic speech recognition systems on low-resource languages. Furthermore, we show that our approach outperforms existing augmentation strategies, offering a practical solution for enhancing speech technology in underrepresented linguistic communities.

📄 PDF Abstract BibTeX arXiv:2511.20534

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionData Augmentation

Similar Papers 제목 키워드 기반

Using Audio Books for Training a Text-to-Speech System

2014-05-01 · LREC 2014 5 · Chalam, Aimilios aris, Pirros Tsiakoulis, Sotiris Karabetsos 외

Creating new voices for a TTS system often requires a costly procedure of designing and recording an audio corpus, a time consuming and effort intensive task. Using publicly available audiobooks as the raw material of a …

DiversitySpeech Synthesistext-to-speechText to Speech

PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

2023-09-17 · Jixun Yao, Yuguang Yang, Yi Lei, Ziqian Ning 외

Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or ref…

Voice Conversion

VoiceAgentBench: Are Voice Assistants ready for agentic tasks?

2025-10-09 · Dhruv Jain, Harshit Shukla, Gautam Rajeev, Ashish Kulkarni 외 arxiv

Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities su…

Adversarial RobustnessQuestion AnsweringVoice Conversion

Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder

2025-04-16 · Soobin Suh, Dabi Ahn, Heewoong Park, Jonghun Park

Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script wi…

DiversityVoice Conversion

VOLTA: Improving Generative Diversity by Variational Mutual Information Maximizing Autoencoder

2023-07-03 · Yueen Ma, Dafeng Chi, Jingjing Li, Kai Song 외

The natural language generation domain has witnessed great success thanks to Transformer models. Although they have achieved state-of-the-art generative quality, they often neglect generative diversity. Prior attempts to…

DiversityText Generation