Relational Data Selection for Data Augmentation of Speaker-dependent Multi-band MelGAN Vocoder
Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available. Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impractical to collect a large amount of data of a specific target speaker for most real-world applications. To tackle the problem of limited target data, a data augmentation method based on speaker representation and similarity measurement of speaker verification is proposed in this paper. The proposed method selects utterances that have similar speaker identity to the target speaker from an external corpus, and then combines the selected utterances with the limited target data for SD vocoder adaptation. The evaluation results show that, compared with the vocoder adapted using only limited target data, the vocoder adapted using augmented data improves both the quality and similarity of synthesized speech.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSpeaker VerificationSimilar Papers 제목 키워드 기반
Graph-Based Feature Augmentation for Predictive Tasks on Relational Datasets
Data has become a foundational asset driving innovation across domains such as finance, healthcare, and e-commerce. In these areas, predictive modeling over relational tables is commonly employed, with increasing emphasi…
Target Speech Extraction Based on Blind Source Separation and X-vector-based Speaker Selection Trained with Data Augmentation
Extracting the desired speech from a mixture is a meaningful and challenging task. The end-to-end DNN-based methods, though attractive, face the problem of generalization. In this paper, we explore a sequential approach …
blind source separationData AugmentationSpeaker RecognitionSpeech ExtractionARDA: Automatic Relational Data Augmentation for Machine Learning
Automatic machine learning (\AML) is a family of techniques to automate the process of training predictive models, aiming to both improve performance and make machine learning more accessible. While many recent works hav…
BIG-bench Machine LearningData Augmentationfeature selectionModel SelectionRelation-aware Graph Attention Networks with Relational Position Encodings for Emotion Recognition in Conversations
Interest in emotion recognition in conversations (ERC) has been increasing in various fields, because it can be used to analyze user behaviors and detect fake news. Many recent ERC methods use graph-based neural networks…
Emotion RecognitionEmotion Recognition in ConversationGraph AttentionPosition+1Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?
Previous work on speaker adaptation for end-to-end speech synthesis still falls short in speaker similarity. We investigate an orthogonal approach to the current speaker adaptation paradigms, speaker augmentation, by cre…
Speech Synthesis