paper-with-me

Papers

HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection

2026-02-01 · Zhili Nicholas Liang, Soyeon Caren Han, Qizhou Wang, Christopher Leckie arxiv

Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich multi-layer representations, existing detectors treat layers independently and overlook temporal and hierarchical dependencies critical for identifying synthetic artefacts. We propose HierCon, a hierarchical layer attention framework combined with margin-based contrastive learning that models dependencies across temporal frames, neighbouring layers, and layer groups, while encouraging domain-invariant embeddings. Evaluated on ASVspoof 2021 DF and In-the-Wild datasets, our method achieves state-of-the-art performance (1.93% and 6.87% EER), improving over independent layer weighting by 36.6% and 22.5% respectively. The results and attention visualisations confirm that hierarchical modelling enhances generalisation to cross-domain generation techniques and recording conditions.

📄 PDF Abstract BibTeX arXiv:2602.01032

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake DetectionContrastive LearningVoice Conversion

Results from the Paper

RankTaskDatasetModelMetrics
#42 Audio Deepfake Detection ASVspoof 2021 HierCon 21DF EER: 1.93

Similar Papers 제목 키워드 기반

Joint Audio-Visual Attention with Contrastive Learning for More General Deepfake Detection

2024-01-22 · ACM Transactions on Multimedia Computing, Communications, and Applications 2024 1 · Yibo Zhang, WEIGUO LIN, andJUNFENG XU

With the continuous advancement of deepfake technology, there has been a surge in the creation of realistic fake videos. Unfortunately, the malicious utilization of deepfake poses a significant threat to societal moralit…

Contrastive LearningDeepFake DetectionFace SwappingHuman Detection of Deepfakes

Do You Really Mean That? Content Driven Audio-Visual Deepfake Dataset and Multimodal Method for Temporal Forgery Localization

2022-04-13 · Zhixi Cai, Kalin Stefanov, Abhinav Dhall, Munawar Hayat

Due to its high societal impact, deepfake detection is getting active attention in the computer vision community. Most deepfake detection methods rely on identity, facial attributes, and adversarial perturbation-based sp…

BenchmarkingDeepFake DetectionTemporal Forgery Localization

Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework

2025-08-04 · Andrea Di Pierno, Luca Guarnera, Dario Allegra, Sebastiano Battiato arxiv

The proliferation of audio deepfakes poses a growing threat to trust in digital communications. While detection methods have advanced, attributing audio deepfakes to their source models remains an underexplored yet cruci…

Audio Deepfake Detection

CLAD: Robust Audio Deepfake Detection Against Manipulation Attacks with Contrastive Learning

2024-04-24 · Haolin Wu, Jing Chen, Ruiying Du, Cong Wu 외

The increasing prevalence of audio deepfakes poses significant security threats, necessitating robust detection methods. While existing detection systems exhibit promise, their robustness against malicious audio manipula…

Audio Deepfake DetectionContrastive LearningDeepFake DetectionFace Swapping

Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization

2024-08-02 · Vinaya Sree Katamneni, Ajita Rattani

In the digital age, the emergence of deepfakes and synthetic media presents a significant threat to societal and political integrity. Deepfakes based on multi-modal manipulation, such as audio-visual, are more realistic …

DeepFake DetectionFace Swapping