HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich multi-layer representations, existing detectors treat layers independently and overlook temporal and hierarchical dependencies critical for identifying synthetic artefacts. We propose HierCon, a hierarchical layer attention framework combined with margin-based contrastive learning that models dependencies across temporal frames, neighbouring layers, and layer groups, while encouraging domain-invariant embeddings. Evaluated on ASVspoof 2021 DF and In-the-Wild datasets, our method achieves state-of-the-art performance (1.93% and 6.87% EER), improving over independent layer weighting by 36.6% and 22.5% respectively. The results and attention visualisations confirm that hierarchical modelling enhances generalisation to cross-domain generation techniques and recording conditions.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio Deepfake DetectionContrastive LearningVoice ConversionResults from the Paper
| Rank | Task | Dataset | Model | Metrics |
|---|---|---|---|---|
| #42 | Audio Deepfake Detection | ASVspoof 2021 | HierCon | 21DF EER: 1.93 |
Similar Papers 제목 키워드 기반
Joint Audio-Visual Attention with Contrastive Learning for More General Deepfake Detection
With the continuous advancement of deepfake technology, there has been a surge in the creation of realistic fake videos. Unfortunately, the malicious utilization of deepfake poses a significant threat to societal moralit…
Contrastive LearningDeepFake DetectionFace SwappingHuman Detection of DeepfakesDo You Really Mean That? Content Driven Audio-Visual Deepfake Dataset and Multimodal Method for Temporal Forgery Localization
Due to its high societal impact, deepfake detection is getting active attention in the computer vision community. Most deepfake detection methods rely on identity, facial attributes, and adversarial perturbation-based sp…
BenchmarkingDeepFake DetectionTemporal Forgery LocalizationTowards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
The proliferation of audio deepfakes poses a growing threat to trust in digital communications. While detection methods have advanced, attributing audio deepfakes to their source models remains an underexplored yet cruci…
Audio Deepfake DetectionCLAD: Robust Audio Deepfake Detection Against Manipulation Attacks with Contrastive Learning
The increasing prevalence of audio deepfakes poses significant security threats, necessitating robust detection methods. While existing detection systems exhibit promise, their robustness against malicious audio manipula…
Audio Deepfake DetectionContrastive LearningDeepFake DetectionFace SwappingContextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
In the digital age, the emergence of deepfakes and synthetic media presents a significant threat to societal and political integrity. Deepfakes based on multi-modal manipulation, such as audio-visual, are more realistic …
DeepFake DetectionFace Swapping