Combining Adversarial Training and Disentangled Speech Representation for Robust Zero-Resource Subword Modeling
This study addresses the problem of unsupervised subword unit discovery from untranscribed speech. It forms the basis of the ultimate goal of ZeroSpeech 2019, building text-to-speech systems without text labels. In this work, unit discovery is formulated as a pipeline of phonetically discriminative feature learning and unit inference. One major difficulty in robust unsupervised feature learning is dealing with speaker variation. Here the robustness towards speaker variation is achieved by applying adversarial training and FHVAE based disentangled speech representation learning. A comparison of the two approaches as well as their combination is studied in a DNN-bottleneck feature (DNN-BNF) architecture. Experiments are conducted on ZeroSpeech 2019 and 2017. Experimental results on ZeroSpeech 2017 show that both approaches are effective while the latter is more prominent, and that their combination brings further marginal improvement in across-speaker condition. Results on ZeroSpeech 2019 show that in the ABX discriminability task, our approaches significantly outperform the official baseline, and are competitive to or even outperform the official topline. The proposed unit sequence smoothing algorithm improves synthesis quality, at a cost of slight decrease in ABX discriminability.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningSpeech Representation Learningtext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Adversarially learning disentangled speech representations for robust multi-factor voice conversion
Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC). Conventional speech representation learning methods in VC only factorize speech a…
Representation LearningRhythmSpeech Representation LearningStyle Transfer+1Talking Face Generation by Adversarially Disentangled Audio-Visual Representation
Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the s…
Face GenerationLip ReadingRetrievalTalking Face Generation+1Powerful Speaker Embedding Training Framework by Adversarially Disentangled Identity Representation
The main challenge of speaker verification in the wild is the interference caused by irrelevant information in speech and the lack of speaker labels in speech datasets. In order to solve the above problems, we propose a …
Speaker VerificationUnsupervised Speech Domain Adaptation Based on Disentangled Representation Learning for Robust Speech Recognition
In general, the performance of automatic speech recognition (ASR) systems is significantly degraded due to the mismatch between training and test environments. Recently, a deep-learning-based image-to-image translation t…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain AdaptationImage-to-Image Translation+5StyleTTS-VC: One-Shot Voice Conversion by Knowledge Transfer from Style-Based TTS Models
One-shot voice conversion (VC) aims to convert speech from any source speaker to an arbitrary target speaker with only a few seconds of reference speech from the target speaker. This relies heavily on disentangling the s…
Data Augmentationtext-to-speechText to SpeechTransfer Learning+1