paper-with-me

Papers

StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation

2026-03-06 · Nikita Kuzmin, Kong Aik Lee, Eng Siong Chng arxiv

We address the challenge of preserving emotional content in streaming speaker anonymization (SA). Neural audio codec language models trained for audio continuation tend to degrade source emotion: content tokens discard emotional information, and the model defaults to dominant acoustic patterns rather than preserving paralinguistic attributes. We propose supervised finetuning with neutral-emotion utterance pairs from the same speaker, combined with frame-level emotion distillation on acoustic token hidden states. All modifications are confined to finetuning, which takes less than 2 hours on 4 GPUs and adds zero inference latency overhead, while maintaining a competitive 180ms streaming latency. On the VoicePrivacy 2024 protocol, our approach achieves a 49.2% UAR (emotion preservation) with 5.77% WER (intelligibility), a +24% relative UAR improvement over the baseline (39.7%->49.2%) and +10% over the emotion-prompt variant (44.6% UAR), while maintaining strong privacy (EER 49.0%). Demo and code are available: https://anonymous3842031239.github.io/

📄 PDF Abstract BibTeX arXiv:2603.06079

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Privacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization

2024-09-05 · Zexin Cai, Henry Li Xinyuan, Ashi Garg, Leibny Paola García-Perera 외

Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while …

Speaker anonymizationSpeaker Verification

Evaluation of Speaker Anonymization on Emotional Speech

2023-04-15 · Hubert Nourtel, Pierre Champion, Denis Jouvet, Anthony Larcher 외

Speech data carries a range of personal information, such as the speaker's identity and emotional state. These attributes can be used for malicious purposes. With the development of virtual assistants, a new generation o…

Automatic Speech RecognitionEmotion RecognitionSpeaker anonymizationspeech-recognition+2

Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models

2026-01-20 · Nikita Kuzmin, Songting Liu, Kong Aik Lee, Eng Siong Chng arxiv

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speak…

Voice Conversion

The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization

2026-01-17 · Natalia Tomashenko, Xiaoxiao Miao, Pierre Champion, Sarina Meyer 외 arxiv

We present results and analyses from the third VoicePrivacy Challenge held in 2024, which focuses on advancing voice anonymization technologies. The task was to develop a voice anonymization system for speech data that c…

End-to-end streaming model for low-latency speech anonymization

2024-06-13 · Waris Quamer, Ricardo Gutierrez-Osuna

Speaker anonymization aims to conceal cues to speaker identity while preserving linguistic content. Current machine learning based approaches require substantial computational resources, hindering real-time streaming app…

DecoderSpeaker anonymization