A Speech Representation Anonymization Framework via Selective Noise Perturbation
Privacy and security are major concerns when communicating speech signals to cloud services such as automatic speech recognition (ASR) and speech emotion recognition (SER). Existing solutions for speech anonymization mainly focus on voice conversion or voice modification to convert a raw utterance into another one with similar content but different, or no, identity-related information. However, an alternative approach to share speech data under the form of privacy-preserving representation has been largely under-explored. In this paper, we propose a speech anonymization framework that achieves privacy via noise perturbation to a selected subset of the high-utility representations extracted using a pre-trained speech encoder. The subset is chosen with a Transformer-based privacy-risk saliency estimator. We validate our framework on four tasks, namely, Automatic Speaker Verification (ASV), ASR, SER and Intent Classification (IC) for privacy and utility assessment. Experimental results show that our approach is able to achieve a competitive, or even better, utility compared to the speech anonymization baselines from the VoicePrivacy2022 Challenges, providing the same level of privacy. Moreover, the easily-controlled amount of perturbation allows our framework to have a flexible range of privacy-utility trade-offs without re-training any component.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognitionintent-classificationIntent ClassificationPrivacy PreservingSpeaker VerificationSpeech Emotion Recognitionspeech-recognitionSpeech RecognitionVoice ConversionSimilar Papers 제목 키워드 기반
Anonymizing Speech: Evaluating and Designing Speaker Anonymization Techniques
The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses s…
QuantizationSpeaker anonymizationVoice CloningVoice ConversionDifferentially Private Speaker Anonymization
Sharing real-world speech utterances is key to the training and deployment of voice-based services. However, it also raises privacy risks as speech contains a wealth of personal data. Speaker anonymization aims to remove…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DisentanglementSpeaker anonymization+2Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech
Self-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-labels, the resulting representations are onl…
Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models
Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization systems are developed on adult speech, leading to degraded performance when appli…
Self-Supervised LearningDomain AdaptationReprogramming Self-supervised Learning-based Speech Representations for Speaker Anonymization
Current speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient…
Self-Supervised LearningSpeaker anonymization