Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy
In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes with a privacy-utility trade-off between protection of individuals and usability of the data for downstream applications. One of the challenges in this context is to create non-existent voices that sound as natural as possible. In this work, we propose to tackle this issue by generating speaker embeddings using a generative adversarial network with Wasserstein distance as cost function. By incorporating these artificial embeddings into a speech-to-text-to-speech pipeline, we outperform previous approaches in terms of privacy and utility. According to standard objective metrics and human evaluation, our approach generates intelligible and content-preserving yet privacy-protecting versions of the original recordings.
Code (1)
Tasks
Generative Adversarial NetworkSpeaker anonymizationSpeech-to-Texttext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Self-Supervised Speech Representations Preserve Speech Characteristics while Anonymizing Voices
Collecting speech data is an important step in training speech recognition systems and other speech-based machine learning models. However, the issue of privacy protection is an increasing concern that must be addressed.…
Speaker Verificationspeech-recognitionSpeech RecognitionVoice ConversionAsynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine…
DisentanglementAnonymizing Speech: Evaluating and Designing Speaker Anonymization Techniques
The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses s…
QuantizationSpeaker anonymizationVoice CloningVoice ConversionSpeaker Identity Preservation in Dysarthric Speech Reconstruction by Adversarial Speaker Adaptation
Dysarthric speech reconstruction (DSR), which aims to improve the quality of dysarthric speech, remains a challenge, not only because we need to restore the speech to be normal, but also must preserve the speaker's ident…
Multi-Task LearningSpeaker VerificationPrivacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization
Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while …
Speaker anonymizationSpeaker Verification