Can we steal your vocal identity from the Internet?: Initial investigation of cloning Obama's voice using GAN, WaveNet and low-quality found data
Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not considered in the ASVspoof2015 database are appearing. Such examples include direct waveform modelling and generative adversarial networks. We also need to investigate the feasibility of training spoofing systems using only low-quality found data. For that purpose, we developed a generative adversarial network-based speech enhancement system that improves the quality of speech data found in publicly available sources. Using the enhanced data, we trained state-of-the-art text-to-speech and voice conversion models and evaluated them in terms of perceptual speech quality and speaker similarity. The results show that the enhancement models significantly improved the SNR of low-quality degraded data found in publicly available sources and that they significantly improved the perceptual cleanliness of the source speech without significantly degrading the naturalness of the voice. However, the results also show limitations when generating speech with the low-quality found data.
Code (0)
등록된 구현이 없습니다.
Tasks
Generative Adversarial NetworkSpeech EnhancementSpeech Synthesistext-to-speechText to SpeechVoice ConversionSimilar Papers 제목 키워드 기반
ANTI-PHISHING IN ANDROID PHONE PROJECT REPORT.
Phishing is a new word produced from 'fishing', it refers to the act that the attacker allure users to visit a faked Web site by sending them faked e-mails (or instant messages), and stealthily get victim's personal …
Turing Test for the Internet of Things
How smart is your kettle? How smart are things in your kitchen, your house, your neighborhood, on the internet? With the advent of Internet of Things, and the move of making devices `smart' by utilizing AI, a natural que…
Singing Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders
We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages re…
DecoderVoice ConversionVocal Style Factorization for Effective Speaker Recognition in Affective Scenarios
The accuracy of automated speaker recognition is negatively impacted by change in emotions in a person's speech. In this paper, we hypothesize that speaker identity is composed of various vocal style factors that may be …
Speaker RecognitionHow Google Search Works
Internet telephony consists of a combination of hardware and software that enables you to use the Internet as the transmission medium for telephone calls. For users who have free, or fixed-price Internet access, Internet…
Form