paper-with-me

Papers

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

2026-04-29 · Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews, Philipp Koehn, Berrak Sisman arxiv

To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing methods either discard prosody for privacy or lack a principled mechanism to control the utility-privacy trade-off, operating at fixed design points. We propose DiffAnon, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation. DiffAnon refines acoustic detail over semantic embeddings of an RVQ codec, enabling smooth interpolation between anonymization strength and prosodic fidelity within a single model. To the best of our knowledge, it is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control. Experiments demonstrate structured trade-off behavior, achieving strong utility while maintaining competitive privacy across controllable operating points.

📄 PDF Abstract BibTeX arXiv:2604.26281

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NTU-NPU System for Voice Privacy 2024 Challenge

2024-10-03 · Nikita Kuzmin, Hieu-Thi Luong, Jixun Yao, Lei Xie 외

In this work, we describe our submissions for the Voice Privacy Challenge 2024. Rather than proposing a novel speech anonymization system, we enhance the provided baselines to meet all required conditions and improve eva…

Disentanglement

Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization

2024-09-24 · Sotheara Leang, Anderson Augusma, Eric Castelli, Frédérique Letué 외

Human speech conveys prosody, linguistic content, and speaker identity. This article investigates a novel speaker anonymization approach using an end-to-end network based on a Vector-Quantized Variational Auto-Encoder (V…

DecoderSpeaker anonymizationSpeaker Identification

UniVoice: A Unified Model for Speech and Singing Voice Generation

2026-06-04 · Junjie Zheng, Huixin Xue, Shihong Ren, Chaofan Ding 외 arxiv

Text-to-speech (TTS) and singing voice synthesis (SVS) both aim to generate human vocal audio from symbolic inputs, but they impose different requirements on the generation process. Speech generation relies on flexible, …

Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion

2025-05-30 · Kaidi Wang, Wenhao Guan, Ziyue Jiang, Hukai Huang 외

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic…

In-Context LearningVoice Conversion

Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses

2024-10-20 · Suhita Ghosh, Tim Thiele, Frederic Lorbeer, Frank Dreyer 외

The increasing use of cloud-based speech assistants has heightened the need for effective speech anonymization, which aims to obscure a speaker's identity while retaining critical information for subsequent tasks. One ap…

Voice Conversion