paper-with-me

홈 › Papers

Are disentangled representations all you need to build speaker anonymization systems?

2022-08-22 · Pierre Champion, Denis Jouvet, Anthony Larcher

Speech signals contain a lot of sensitive information, such as the speaker's identity, which raises privacy concerns when speech data get collected. Speaker anonymization aims to transform a speech signal to remove the source speaker's identity while leaving the spoken content unchanged. Current methods perform the transformation by relying on content/speaker disentanglement and voice conversion. Usually, an acoustic model from an automatic speech recognition system extracts the content representation while an x-vector system extracts the speaker representation. Prior work has shown that the extracted features are not perfectly disentangled. This paper tackles how to improve features disentanglement, and thus the converted anonymized speech. We propose enhancing the disentanglement by removing speaker information from the acoustic model using vector quantization. Evaluation done using the VoicePrivacy 2022 toolkit showed that vector quantization helps conceal the original speaker identity while maintaining utility for speech recognition.

📄 PDF Abstract BibTeX arXiv:2208.10497

Code (0)

등록된 구현이 없습니다.

Tasks

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)DisentanglementQuantizationSpeaker anonymizationspeech-recognitionSpeech RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

End-to-end streaming model for low-latency speech anonymization

2024-06-13 · Waris Quamer, Ricardo Gutierrez-Osuna

Speaker anonymization aims to conceal cues to speaker identity while preserving linguistic content. Current machine learning based approaches require substantial computational resources, hindering real-time streaming app…

DecoderSpeaker anonymization

NPU-NTU System for Voice Privacy 2024 Challenge

2024-09-06 · Jixun Yao, Nikita Kuzmin, Qing Wang, Pengcheng Guo 외

Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair be…

DisentanglementSpeaker anonymization

Reprogramming Self-supervised Learning-based Speech Representations for Speaker Anonymization

2023-11-17 · Xiaojiao Chen, Sheng Li, Jiyi Li, Hao Huang 외

Current speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient…

Self-Supervised LearningSpeaker anonymization

DAST: A Dual-Stream Voice Anonymization Attacker with Staged Training

2026-03-13 · Ridwan Arefeen, Xiaoxiao Miao, Rong Tong, Aik Beng Ng 외 arxiv

Voice anonymization masks vocal traits while preserving linguistic content, which may still leak speaker-specific patterns. To assess and strengthen privacy evaluation, we propose a dual-stream attacker that fuses spectr…

Self-Supervised LearningVoice Conversion

Target speaker anonymization in multi-speaker recordings

2025-10-10 · Natalia Tomashenko, Junichi Yamagishi, Xin Wang, Yun Liu 외 arxiv

Most of the existing speaker anonymization research has focused on single-speaker audio, leading to the development of techniques and evaluation metrics optimized for such condition. This study addresses the significant …