paper-with-me

Papers

Speaker anonymization using neural audio codec language models

2023-09-25 · Michele Panariello, Francesco Nespoli, Massimiliano Todisco, Nicholas Evans

The vast majority of approaches to speaker anonymization involve the extraction of fundamental frequency estimates, linguistic features and a speaker embedding which is perturbed to obfuscate the speaker identity before an anonymized speech waveform is resynthesized using a vocoder. Recent work has shown that x-vector transformations are difficult to control consistently: other sources of speaker information contained within fundamental frequency and linguistic features are re-entangled upon vocoding, meaning that anonymized speech signals still contain speaker information. We propose an approach based upon neural audio codecs (NACs), which are known to generate high-quality synthetic speech when combined with language models. NACs use quantized codes, which are known to effectively bottleneck speaker-related information: we demonstrate the potential of speaker anonymization systems based on NAC language modeling by applying the evaluation framework of the Voice Privacy Challenge 2022.

📄 PDF Abstract BibTeX arXiv:2309.14129

Code (3)

eurecom-asp/spk_anon_nac_lm 공식 구현 pytorch
m-pana/spk_anon_nac_lm pytorch
voice-privacy-challenge/voice-privacy-challenge-2024 pytorch

Tasks

Language ModelingLanguage ModellingSpeaker anonymization

Similar Papers 제목 키워드 기반

Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models

2026-01-20 · Nikita Kuzmin, Songting Liu, Kong Aik Lee, Eng Siong Chng arxiv

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speak…

Voice Conversion

StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation

2026-03-06 · Nikita Kuzmin, Kong Aik Lee, Eng Siong Chng arxiv

We address the challenge of preserving emotional content in streaming speaker anonymization (SA). Neural audio codec language models trained for audio continuation tend to degrade source emotion: content tokens discard e…

Target speaker anonymization in multi-speaker recordings

2025-10-10 · Natalia Tomashenko, Junichi Yamagishi, Xin Wang, Yun Liu 외 arxiv

Most of the existing speaker anonymization research has focused on single-speaker audio, leading to the development of techniques and evaluation metrics optimized for such condition. This study addresses the significant …

Universal Semantic Disentangled Privacy-preserving Speech Representation Learning

2025-05-19 · Biel Tura Vecino, Subhadeep Maji, Aravind Varier, Antonio Bonafonte 외

The use of audio recordings of human speech to train LLMs poses privacy concerns due to these models' potential to generate outputs that closely resemble artifacts in the training data. In this study, we propose a speake…

DecoderPrivacy PreservingRepresentation LearningSpeaker anonymization+1

NPU-NTU System for Voice Privacy 2024 Challenge

2024-09-06 · Jixun Yao, Nikita Kuzmin, Qing Wang, Pengcheng Guo 외

Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair be…

DisentanglementSpeaker anonymization