paper-with-me

Papers

End-to-end streaming model for low-latency speech anonymization

2024-06-13 · Waris Quamer, Ricardo Gutierrez-Osuna

Speaker anonymization aims to conceal cues to speaker identity while preserving linguistic content. Current machine learning based approaches require substantial computational resources, hindering real-time streaming applications. To address these concerns, we propose a streaming model that achieves speaker anonymization with low latency. The system is trained in an end-to-end autoencoder fashion using a lightweight content encoder that extracts HuBERT-like information, a pretrained speaker encoder that extract speaker identity, and a variance encoder that injects pitch and energy information. These three disentangled representations are fed to a decoder that re-synthesizes the speech signal. We present evaluation results from two implementations of our system, a full model that achieves a latency of 230ms, and a lite version (0.1x in size) that further reduces latency to 66ms while maintaining state-of-the-art performance in naturalness, intelligibility, and privacy preservation.

📄 PDF Abstract BibTeX arXiv:2406.09277

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeaker anonymization

Similar Papers 제목 키워드 기반

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

2026-02-10 · Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna arxiv

Real-time voice conversion and speaker anonymization require causal, low-latency synthesis without sacrificing intelligibility or naturalness. Current systems have a core representational mismatch: content is time-varyin…

Voice ConversionSpeech Synthesis

DarkStream: real-time speech anonymization with low latency

2025-09-04 · Waris Quamer, Ricardo Gutierrez-Osuna arxiv

We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahea…

Speaker VerificationSpeech Synthesis

StreamVC: Real-Time Low-Latency Voice Conversion

2024-01-05 · Yang Yang, Yury Kartynnik, Yunpeng Li, Jiuqiang Tang 외

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces…

Speech SynthesisVoice Conversion

Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models

2026-01-20 · Nikita Kuzmin, Songting Liu, Kong Aik Lee, Eng Siong Chng arxiv

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speak…

Voice Conversion

StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation

2026-03-06 · Nikita Kuzmin, Kong Aik Lee, Eng Siong Chng arxiv

We address the challenge of preserving emotional content in streaming speaker anonymization (SA). Neural audio codec language models trained for audio continuation tend to degrade source emotion: content tokens discard e…