paper-with-me

홈 › Papers

Mitigating Unauthorized Speech Synthesis for Voice Protection

2024-10-28 · Zhisheng Zhang, Qianyi Yang, Derui Wang, Pengyang Huang, Yuxin Cao, Kai Ye, Jie Hao

With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our daily lives. Therefore, it is crucial to protect publicly accessible speech data that contains sensitive information, such as personal voiceprints. Most previous defense methods have focused on spoofing speaker verification systems in timbre similarity but the synthesized deepfake speech is still of high quality. In response to the rising hazards, we devise an effective, transferable, and robust proactive protection technology named Pivotal Objective Perturbation (POP) that applies imperceptible error-minimizing noises on original speech samples to prevent them from being effectively learned for text-to-speech (TTS) synthesis models so that high-quality deepfake speeches cannot be generated. We conduct extensive experiments on state-of-the-art (SOTA) TTS models utilizing objective and subjective metrics to comprehensively evaluate our proposed method. The experimental results demonstrate outstanding effectiveness and transferability across various models. Compared to the speech unclarity score of 21.94% from voice synthesizers trained on samples without protection, POP-protected samples significantly increase it to 127.31%. Moreover, our method shows robustness against noise reduction and data augmentation techniques, thereby greatly reducing potential hazards.

📄 PDF Abstract BibTeX arXiv:2410.20742

Code (1)

wxzyd123/pivotal_objective_perturbation 공식 구현 pytorch

Tasks

Data AugmentationFace SwappingSpeaker VerificationSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations

2025-05-19 · Seungmin Kim, Sohee Park, Donghyun Kim, Jisu Lee 외

With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices…

Speaker VerificationSpeech EnhancementSpeech Synthesis

SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis

2025-04-14 · Zhisheng Zhang, Derui Wang, Qianyi Yang, Pengyang Huang 외

Speech synthesis technology has brought great convenience, while the widespread usage of realistic deepfake audio has triggered hazards. Malicious adversaries may unauthorizedly collect victims' speeches and clone a simi…

Face SwappingSpeech Synthesis

VocalBridge: Latent Diffusion-Bridge Purification for Defeating Perturbation-Based Voiceprint Defenses

2026-01-05 · Maryam Abbasihafshejani, AHM Nazmus Sakib, Murtuza Jadliwala arxiv

The rapid advancement of speech synthesis technologies, including text-to-speech (TTS) and voice conversion (VC), has intensified security and privacy concerns related to voice cloning. Recent defenses attempt to prevent…

Speaker VerificationSpeech RecognitionVoice ConversionSpeech Synthesis

SingFake: Singing Voice Deepfake Detection

2023-09-14 · Yongyi Zang, You Zhang, Mojtaba Heydari, Zhiyao Duan

The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage. Unlike synthesized speech, synthesized singing voices are typically released in songs c…

DeepFake DetectionFace SwappingSinging Voice SynthesisSynthetic Speech Detection

"Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World

2021-09-20 · Emily Wenger, Max Bronckers, Christian Cianfarani, Jenna Cryan 외

Advances in deep learning have introduced a new wave of voice synthesis tools, capable of producing audio that sounds as if spoken by a target speaker. If successful, such tools in the wrong hands will enable a range of …

Deep LearningSpeaker RecognitionSpeech Synthesis