paper-with-me

홈 › Papers

RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations

2025-05-19 · Seungmin Kim, Sohee Park, Donghyun Kim, Jisu Lee, Daeseon Choi

With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices. Existing defenses that inject adversarial perturbations directly into audio signals have limited effectiveness, as these perturbations can easily be neutralized by speech enhancement methods. To overcome this limitation, we propose RoVo (Robust Voice), a novel proactive defense technique that injects adversarial perturbations into high-dimensional embedding vectors of audio signals, reconstructing them into protected speech. This approach effectively defends against speech synthesis attacks and also provides strong resistance to speech enhancement models, which represent a secondary attack threat. In extensive experiments, RoVo increased the Defense Success Rate (DSR) by over 70% compared to unprotected speech, across four state-of-the-art speech synthesis models. Specifically, RoVo achieved a DSR of 99.5% on a commercial speaker-verification API, effectively neutralizing speech synthesis attack. Moreover, RoVo's perturbations remained robust even under strong speech enhancement conditions, outperforming traditional methods. A user study confirmed that RoVo preserves both naturalness and usability of protected speech, highlighting its effectiveness in complex and evolving threat scenarios.

📄 PDF Abstract BibTeX arXiv:2505.12686

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationSpeech EnhancementSpeech Synthesis

Similar Papers 제목 키워드 기반

SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis

2025-04-14 · Zhisheng Zhang, Derui Wang, Qianyi Yang, Pengyang Huang 외

Speech synthesis technology has brought great convenience, while the widespread usage of realistic deepfake audio has triggered hazards. Malicious adversaries may unauthorizedly collect victims' speeches and clone a simi…

Face SwappingSpeech Synthesis

Mitigating Unauthorized Speech Synthesis for Voice Protection

2024-10-28 · Zhisheng Zhang, Qianyi Yang, Derui Wang, Pengyang Huang 외

With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our…

Data AugmentationFace SwappingSpeaker VerificationSpeech Synthesis+2

Privacy against Real-Time Speech Emotion Detection via Acoustic Adversarial Evasion of Machine Learning

2022-11-17 · Brian Testa, Yi Xiao, Harshit Sharma, Avery Gump 외

Smart speaker voice assistants (VAs) such as Amazon Echo and Google Home have been widely adopted due to their seamless integration with smart home devices and the Internet of Things (IoT) technologies. These VA services…

Emotion RecognitionSpeech Emotion Recognition

VocalBridge: Latent Diffusion-Bridge Purification for Defeating Perturbation-Based Voiceprint Defenses

2026-01-05 · Maryam Abbasihafshejani, AHM Nazmus Sakib, Murtuza Jadliwala arxiv

The rapid advancement of speech synthesis technologies, including text-to-speech (TTS) and voice conversion (VC), has intensified security and privacy concerns related to voice cloning. Recent defenses attempt to prevent…

Speaker VerificationSpeech RecognitionVoice ConversionSpeech Synthesis

VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning

2025-05-18 · Qianyue Hu, Junyan Wu, Wei Lu, Xiangyang Luo

Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrup…

Representation LearningVoice Cloning