paper-with-me

Papers

DarkStream: real-time speech anonymization with low latency

2025-09-04 · Waris Quamer, Ricardo Gutierrez-Osuna arxiv

We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahead buffer, and transformer-based contextual layers. To further reduce inference time, the model generates waveforms directly via a neural vocoder, thus removing intermediate mel-spectrogram conversions. Finally, DarkStream anonymizes speaker identity by injecting a GAN-generated pseudo-speaker embedding into linguistic features from the content encoder. Evaluations show our model achieves strong anonymization, yielding close to 50% speaker verification EER (near-chance performance) on the lazy-informed attack scenario, while maintaining acceptable linguistic intelligibility (WER within 9%). By balancing low-latency, robust privacy, and minimal intelligibility degradation, DarkStream provides a practical solution for privacy-preserving real-time speech communication.

📄 PDF Abstract BibTeX arXiv:2509.04667

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationSpeech Synthesis

Similar Papers 제목 키워드 기반

Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models

2026-01-20 · Nikita Kuzmin, Songting Liu, Kong Aik Lee, Eng Siong Chng arxiv

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speak…

Voice Conversion

End-to-end streaming model for low-latency speech anonymization

2024-06-13 · Waris Quamer, Ricardo Gutierrez-Osuna

Speaker anonymization aims to conceal cues to speaker identity while preserving linguistic content. Current machine learning based approaches require substantial computational resources, hindering real-time streaming app…

DecoderSpeaker anonymization

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

2026-02-10 · Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna arxiv

Real-time voice conversion and speaker anonymization require causal, low-latency synthesis without sacrificing intelligibility or naturalness. Current systems have a core representational mismatch: content is time-varyin…

Voice ConversionSpeech Synthesis

StreamVC: Real-Time Low-Latency Voice Conversion

2024-01-05 · Yang Yang, Yury Kartynnik, Yunpeng Li, Jiuqiang Tang 외

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces…

Speech SynthesisVoice Conversion

SAIC: Integration of Speech Anonymization and Identity Classification

2023-12-23 · Ming Cheng, Xingjian Diao, Shitong Cheng, Wenjun Liu

Speech anonymization and de-identification have garnered significant attention recently, especially in the healthcare area including telehealth consultations, patient voiceprint matching, and patient real-time monitoring…

ClassificationDe-identification