Why disentanglement-based speaker anonymization systems fail at preserving emotions?
Disentanglement-based speaker anonymization involves decomposing speech into a semantically meaningful representation, altering the speaker embedding, and resynthesizing a waveform using a neural vocoder. State-of-the-art systems of this kind are known to remove emotion information. Possible reasons include mode collapse in GAN-based vocoders, unintended modeling and modification of emotions through speaker embeddings, or excessive sanitization of the intermediate representation. In this paper, we conduct a comprehensive evaluation of a state-of-the-art speaker anonymization system to understand the underlying causes. We conclude that the main reason is the lack of emotion-related information in the intermediate representation. The speaker embeddings also have a high impact, if they are learned in a generative context. The vocoder's out-of-distribution performance has a smaller impact. Additionally, we discovered that synthesis artifacts increase spectral kurtosis, biasing emotion recognition evaluation towards classifying utterances as angry. Therefore, we conclude that reporting unweighted average recall alone for emotion recognition performance is suboptimal.
Code (0)
등록된 구현이 없습니다.
Tasks
DisentanglementEmotion RecognitionSpeaker anonymizationSimilar Papers 제목 키워드 기반
MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information of the original speech. While most prio…
DisentanglementSpeaker anonymizationEASY: Emotion-aware Speaker Anonymization via Factorized Distillation
Emotion plays a significant role in speech interaction, conveyed through tone, pitch, and rhythm, enabling the expression of feelings and intentions beyond words to create a more personalized experience. However, most ex…
AttributeDisentanglementRhythmSpeaker anonymizationNPU-NTU System for Voice Privacy 2024 Challenge
Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair be…
DisentanglementSpeaker anonymizationA Benchmark for Multi-speaker Anonymization
Privacy-preserving voice protection approaches primarily suppress privacy-related information derived from paralinguistic attributes while preserving the linguistic content. Existing solutions focus particularly on singl…
BenchmarkingDisentanglementPrivacy PreservingSpeaker anonymization+2Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine…
Disentanglement