paper-with-me

홈 › Papers

Echoes: A semantically-aligned music deepfake detection dataset

2026-03-24 · Octavian Pascu, Dan Oneata, Horia Cucu, Nicolas M. Muller arxiv

We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 4,468 tracks (131 hours of audio) spanning multiple genres (pop, rock, electronic), and includes content generated by ten popular AI music generation systems. To prevent shortcut learning and promote robust generalization, the dataset is deliberately constructed to be challenging, enforcing semantic-level alignment between spoofed audio and bona fide references. This alignment is achieved by conditioning generated audio samples directly on bona-fide waveforms or song descriptors. We evaluate Echoes in a cross-dataset setting against three existing AI-generated music datasets using state-of-the-art Wav2Vec2 XLS-R 2B representations. Results show that (i) Echoes is the hardest in-domain dataset; (ii) detectors trained on existing datasets transfer poorly to Echoes; (iii) training on Echoes yields the strongest generalization performance. These findings suggest that provider diversity and semantic alignment help learn more transferable detection cues.

📄 PDF Abstract BibTeX arXiv:2603.23667

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake DetectionMusic Generation

Similar Papers 제목 키워드 기반

SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan

2024-05-08 · You Zhang, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto 외

The rapid advancement of AI-generated singing voices, which now closely mimic natural human singing and align seamlessly with musical scores, has led to heightened concerns for artists and the music industry. Unlike spok…

DeepFake DetectionFace Swapping

Training-Free Multimodal Deepfake Detection via Graph Reasoning

2025-09-26 · Yuxin Liu, Fei Wang, Kun Li, Yiqi Nie 외 arxiv

Multimodal deepfake detection (MDD) aims to uncover manipulations across visual, textual, and auditory modalities, thereby reinforcing the reliability of modern information systems. Although large vision-language models …

Multimodal ReasoningDeepFake Detection

Detecting Musical Deepfakes

2025-05-03 · Nick Sunday

The proliferation of Text-to-Music (TTM) platforms has democratized music creation, enabling users to effortlessly generate high-quality compositions. However, this innovation also presents new challenges to musicians an…

Face Swapping

Detecting music deepfakes is easy but actually hard

2024-05-07 · Darius Afchar, Gabriel Meseguer-Brocal, Romain Hennequin

In the face of a new era of generative models, the detection of artificially generated content has become a matter of utmost importance. The ability to create credible minute-long music deepfakes in a few seconds on user…

DeepFake DetectionFace Swapping

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

2026-04-09 · Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo 외 arxiv

The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. …

Audio Deepfake Detection