paper-with-me

홈 › Papers

Model Attribution and Detection of Synthetic Speech via Vocoder Fingerprints

2024-11-21 · Matías Pizarro, Mike Laszkiewicz, Shawkat Hesso, Dorothea Kolossa, Asja Fischer

As speech generation technology advances, so do the potential threats of misusing synthetic speech signals. This work tackles three tasks: (1) single-model attribution in an open-world setting corresponding to the task of identifying whether synthetic speech signals originate from a specific vocoder (which requires only target vocoder data), (2) model attribution in a closed-world setting that corresponds to selecting the specific model that generated a sample from a given set of models, and (3) distinguishing synthetic from real speech. We show that standardized average residuals between audio signals and their low-pass or EnCodec filtered versions serve as powerful vocoder fingerprints that can be leveraged for all tasks achieving an average AUROC of over 99% on LJSpeech and JSUT in most settings. The accompanying robustness study shows that it is also resilient to noise levels up to a certain degree.

📄 PDF Abstract BibTeX arXiv:2411.14013

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

An Initial Investigation for Detecting Vocoder Fingerprints of Fake Audio

2022-08-20 · Xinrui Yan, Jiangyan Yi, JianHua Tao, Chenglong Wang 외

Many effective attempts have been made for fake audio detection. However, they can only provide detection results but no countermeasures to curb this harm. For many related practical applications, what model or algorithm…

STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution

2025-05-26 · Anton Firc, Manasi Chibber, Jagabandhu Mishra, Vishwanath Pratap Singh 외

A key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generati…

DeepFake DetectionFace Swapping

An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization

2024-09-17 · Manasi Chhibber, Jagabandhu Mishra, Hyejin Shim, Tomi H. Kinnunen

We propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose d…

Attribute

ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

2025-02-08 · Ashi Garg, Zexin Cai, Lin Zhang, Henry Li Xinyuan 외

The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what extent can low error rates on academic b…

BenchmarkingSelf-Supervised LearningSynthetic Speech Detectiontext-to-speech+1

Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints

2026-04-29 · Wasim Ahmad, Wei Zhang, Xuerui Mao arxiv

Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection has shown promise, most approaches are bin…

Binary ClassificationDeepFake Detection