paper-with-me

홈 › Papers

Synth2Aug: Cross-domain speaker recognition with TTS synthesized speech

2020-11-24 · Yiling Huang, Yutian Chen, Jason Pelecanos, Quan Wang

In recent years, Text-To-Speech (TTS) has been used as a data augmentation technique for speech recognition to help complement inadequacies in the training data. Correspondingly, we investigate the use of a multi-speaker TTS system to synthesize speech in support of speaker recognition. In this study we focus the analysis on tasks where a relatively small number of speakers is available for training. We observe on our datasets that TTS synthesized speech improves cross-domain speaker recognition performance and can be combined effectively with multi-style training. Additionally, we explore the effectiveness of different types of text transcripts used for TTS synthesis. Results suggest that matching the textual content of the target domain is a good practice, and if that is not feasible, a transcript with a sufficiently large vocabulary is recommended.

📄 PDF Abstract BibTeX arXiv:2011.11818

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationSpeaker Recognitionspeech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Speech Recognition with Augmented Synthesized Speech

2019-09-25 · Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia 외

Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribe…

Data AugmentationDiversityRobust Speech Recognitionspeech-recognition+2

An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS

2025-06-25 · Marie Kunešová, Zdeněk Hanzlíček, Jindřich Matoušek

Zero-shot multi-speaker text-to-speech (TTS) systems rely on speaker embeddings to synthesize speech in the voice of an unseen speaker, using only a short reference utterance. While many speaker embeddings have been deve…

Speaker Recognitiontext-to-speechText to SpeechZero-Shot Multi-Speaker TTS

ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion

2022-03-29 · Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Candido Junior 외

We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied to data augmentation for automatic speech recognition (ASR) systems in low/medium-resource scenarios. Through extensive e…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3

What does a network layer hear? Analyzing hidden representations of end-to-end ASR through speech synthesis

2019-11-04 · Chung-Yi Li, Pei-Chieh Yuan, Hung-Yi Lee

End-to-end speech recognition systems have achieved competitive results compared to traditional systems. However, the complex transformations involved between layers given highly variable acoustic signals are hard to ana…

Speaker VerificationSpeech Enhancementspeech-recognitionSpeech Recognition+1

Efficient ASR Training with Conversations that Never Happened

2026-06-02 · Máté Gedeon, Péter Mihajlik arxiv

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with…

Speech Recognition