Vocoder drift compensation by x-vector alignment in speaker anonymisation
For the most popular x-vector-based approaches to speaker anonymisation, the bulk of the anonymisation can stem from vocoding rather than from the core anonymisation function which is used to substitute an original speaker x-vector with that of a fictitious pseudo-speaker. This phenomenon can impede the design of better anonymisation systems since there is a lack of fine-grained control over the x-vector space. The work reported in this paper explores the origin of so-called vocoder drift and shows that it is due to the mismatch between the substituted x-vector and the original representations of the linguistic content, intonation and prosody. Also reported is an original approach to vocoder drift compensation. While anonymisation performance degrades as expected, compensation reduces vocoder drift substantially, offers improved control over the x-vector space and lays a foundation for the design of better anonymisation functions in the future.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Vocoder drift in x-vector-based speaker anonymization
State-of-the-art approaches to speaker anonymization typically employ some form of perturbation function to conceal speaker information contained within an x-vector embedding, then resynthesize utterances in the voice of…
Speaker anonymizationRelational Data Selection for Data Augmentation of Speaker-dependent Multi-band MelGAN Vocoder
Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available. Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impract…
Data AugmentationSpeaker VerificationSpeaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement
Traditional speech enhancement systems produce speech with compromised quality. Here we propose to use the high quality speech generation capability of neural vocoders for better quality speech enhancement. We term this …
ResynthesisSpeech EnhancementData augmentation versus noise compensation for x- vector speaker recognition systems in noisy environments
The explosion of available speech data and new speaker modeling methods based on deep neural networks (DNN) have given the ability to develop more robust speaker recognition systems. Among DNN speaker modelling technique…
Data AugmentationDenoisingSpeaker RecognitionAdaVocoder: Adaptive Vocoder for Custom Voice
Custom voice is to construct a personal speech synthesis system by adapting the source speech synthesis model to the target model through the target few recordings. The solution to constructing a custom voice is to combi…
Speech SynthesisTransfer Learning