Speaker Anonymization with Phonetic Intermediate Representations
In this work, we propose a speaker anonymization pipeline that leverages high quality automatic speech recognition and synthesis systems to generate speech conditioned on phonetic transcriptions and anonymized speaker embeddings. Using phones as the intermediate representation ensures near complete elimination of speaker identity information from the input while preserving the original phonetic content as much as possible. Our experimental results on LibriSpeech and VCTK corpora reveal two key findings: 1) although automatic speech recognition produces imperfect transcriptions, our neural speech synthesis system can handle such errors, making our system feasible and robust, and 2) combining speaker embeddings from different resources is beneficial and their appropriate normalization is crucial. Overall, our final best system outperforms significantly the baselines provided in the Voice Privacy Challenge 2020 in terms of privacy robustness against a lazy-informed attacker while maintaining high intelligibility and naturalness of the anonymized speech.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker anonymizationspeech-recognitionSpeech RecognitionSpeech SynthesisSimilar Papers 제목 키워드 기반
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verific…
Speaker VerificationNWPU-ASLP System for the VoicePrivacy 2022 Challenge
This paper presents the NWPU-ASLP speaker anonymization system for VoicePrivacy 2022 Challenge. Our submission does not involve additional Automatic Speaker Verification (ASV) model or x-vector pool. Our system consists …
Speaker anonymizationSpeaker VerificationWhy disentanglement-based speaker anonymization systems fail at preserving emotions?
Disentanglement-based speaker anonymization involves decomposing speech into a semantically meaningful representation, altering the speaker embedding, and resynthesizing a waveform using a neural vocoder. State-of-the-ar…
DisentanglementEmotion RecognitionSpeaker anonymizationReprogramming Self-supervised Learning-based Speech Representations for Speaker Anonymization
Current speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient…
Self-Supervised LearningSpeaker anonymizationAutomatic Voice Identification after Speech Resynthesis using PPG
Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, v…
ResynthesisSpeaker VerificationVoice Conversion