paper-with-me

홈 › Papers

Deep Learning-based F0 Synthesis for Speaker Anonymization

2023-06-29 · Ünal Ege Gaznepoglu, Nils Peters

Voice conversion for speaker anonymization is an emerging concept for privacy protection. In a deep learning setting, this is achieved by extracting multiple features from speech, altering the speaker identity, and waveform synthesis. However, many existing systems do not modify fundamental frequency (F0) trajectories, which convey prosody information and can reveal speaker identity. Moreover, mismatch between F0 and other features can degrade speech quality and intelligibility. In this paper, we formally introduce a method that synthesizes F0 trajectories from other speech features and evaluate its reconstructional capabilities. Then we test our approach within a speaker anonymization framework, comparing it to a baseline and a state-of-the-art F0 modification that utilizes speaker information. The results show that our method improves both speaker anonymity, measured by the equal error rate, and utility, measured by the word error rate.

📄 PDF Abstract BibTeX arXiv:2306.16860

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningSpeaker anonymizationVoice Conversion

Similar Papers 제목 키워드 기반

Exploring the Importance of F0 Trajectories for Speaker Anonymization using X-vectors and Neural Waveform Models

2021-10-13 · Ünal Ege Gaznepoglu, Nils Peters

Voice conversion for speaker anonymization is an emerging field in speech processing research. Many state-of-the-art approaches are based on the resynthesis of the phoneme posteriorgrams (PPG), the fundamental frequency …

ResynthesisSpeaker anonymizationVoice Conversion

Probing the Feasibility of Multilingual Speaker Anonymization

2024-07-03 · Sarina Meyer, Florian Lux, Ngoc Thang Vu

In speaker anonymization, speech recordings are modified in a way that the identity of the speaker remains hidden. While this technology could help to protect the privacy of individuals around the globe, current research…

Speaker anonymizationSpeech Synthesis

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

2026-02-10 · Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna arxiv

Real-time voice conversion and speaker anonymization require causal, low-latency synthesis without sacrificing intelligibility or naturalness. Current systems have a core representational mismatch: content is time-varyin…

Voice ConversionSpeech Synthesis

DarkStream: real-time speech anonymization with low latency

2025-09-04 · Waris Quamer, Ricardo Gutierrez-Osuna arxiv

We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahea…

Speaker VerificationSpeech Synthesis

Speaker Anonymization with Phonetic Intermediate Representations

2022-07-11 · Sarina Meyer, Florian Lux, Pavel Denisov, Julia Koch 외

In this work, we propose a speaker anonymization pipeline that leverages high quality automatic speech recognition and synthesis systems to generate speech conditioned on phonetic transcriptions and anonymized speaker em…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker anonymizationspeech-recognition+2