paper-with-me

홈 › Papers

SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis

2025-04-14 · Zhisheng Zhang, Derui Wang, Qianyi Yang, Pengyang Huang, Junhan Pu, Yuxin Cao, Kai Ye, Jie Hao, Yixian Yang

Speech synthesis technology has brought great convenience, while the widespread usage of realistic deepfake audio has triggered hazards. Malicious adversaries may unauthorizedly collect victims' speeches and clone a similar voice for illegal exploitation (\textit{e.g.}, telecom fraud). However, the existing defense methods cannot effectively prevent deepfake exploitation and are vulnerable to robust training techniques. Therefore, a more effective and robust data protection method is urgently needed. In response, we propose a defensive framework, \textit{\textbf{SafeSpeech}}, which protects the users' audio before uploading by embedding imperceptible perturbations on original speeches to prevent high-quality synthetic speech. In SafeSpeech, we devise a robust and universal proactive protection technique, \textbf{S}peech \textbf{PE}rturbative \textbf{C}oncealment (\textbf{SPEC}), that leverages a surrogate model to generate universally applicable perturbation for generative synthetic models. Moreover, we optimize the human perception of embedded perturbation in terms of time and frequency domains. To evaluate our method comprehensively, we conduct extensive experiments across advanced models and datasets, both subjectively and objectively. Our experimental results demonstrate that SafeSpeech achieves state-of-the-art (SOTA) voice protection effectiveness and transferability and is highly robust against advanced adaptive adversaries. Moreover, SafeSpeech has real-time capability in real-world tests. The source code is available at \href{https://github.com/wxzyd123/SafeSpeech}{https://github.com/wxzyd123/SafeSpeech}.

📄 PDF Abstract BibTeX arXiv:2504.09839

Code (1)

wxzyd123/safespeech 공식 구현 pytorch

Tasks

Face SwappingSpeech Synthesis

Similar Papers 제목 키워드 기반

Evaluating Voice Conversion-based Privacy Protection against Informed Attackers

2019-11-10 · Brij Mohan Lal Srivastava, Nathalie Vauquier, Md Sahidullah, Aurélien Bellet 외

Speech data conveys sensitive speaker attributes like identity or accent. With a small amount of found data, such attributes can be inferred and exploited for malicious purposes: voice cloning, spoofing, etc. Anonymizati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+3

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning

2025-05-25 · Renyuan Li, Zhibo Liang, Haichuan Zhang, Tianyu Shi 외

Recent breakthroughs in text-to-speech (TTS) voice cloning have raised serious privacy concerns, allowing highly accurate vocal identity replication from just a few seconds of reference audio, while retaining the speaker…

text-to-speechText to SpeechVoice Cloning

Mitigating Unauthorized Speech Synthesis for Voice Protection

2024-10-28 · Zhisheng Zhang, Qianyi Yang, Derui Wang, Pengyang Huang 외

With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our…

Data AugmentationFace SwappingSpeaker VerificationSpeech Synthesis+2

Universal Adversarial Head: Practical Protection against Video Data Leakage

2021-06-18 · ICML Workshop AML 2021 7 · Jiawang Bai, Bin Chen, Dongxian Wu, Chaoning Zhang 외

While online video sharing becomes more popular, it also causes unconscious leakage of personal information in the video retrieval systems like deep hashing. An adversary can collect users' private information from the v…

Deep HashingRetrievalVideo Retrieval

Adversarial speech for voice privacy protection from Personalized Speech generation

2024-01-22 · Shihao Chen, Liping Chen, Jie Zhang, KongAik Lee 외

The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human list…

Speaker Verificationtext-to-speechText to SpeechVoice Conversion