paper-with-me

Papers

Towards Generalized Source Tracing for Codec-Based Deepfake Speech

2025-06-08 · Xuanjun Chen, I-Ming Lin, Lin Zhang, Haibin Wu, Hung-Yi Lee, Jyh-Shing Roger Jang

Recent attempts at source tracing for codec-based deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-the-art performance on the CoSG test set of the CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing.

📄 PDF Abstract BibTeX arXiv:2506.07294

Code (0)

등록된 구현이 없습니다.

Tasks

Face Swapping

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition

2025-01-11 · Yuankun Xie, Xiaopeng Wang, Zhiyong Wang, Ruibo Fu 외

Current research in audio deepfake detection is gradually transitioning from binary classification to multi-class tasks, referred as audio deepfake source tracing task. However, existing studies on source tracing conside…

Audio Deepfake DetectionBinary ClassificationDeepFake DetectionFace Swapping

Multilingual Source Tracing of Speech Deepfakes: A First Benchmark

2025-08-06 · Xi Xuan, Yang Xiao, Rohan Kumar Das, Tomi Kinnunen arxiv

Recent progress in generative AI has made it increasingly easy to create natural-sounding deepfake speech from just a few seconds of audio. While these tools support helpful applications, they also raise serious concerns…

The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio

2024-05-08 · Yuankun Xie, Yi Lu, Ruibo Fu, Zhengqi Wen 외

With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for generalized detection methods. ALM-based deepfake audio currently exhibits widespread, high deception, and type versa…

Audio Deepfake DetectionAudio GenerationDeepFake DetectionFace Swapping+2

CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems

2024-06-11 · Haibin Wu, Yuan Tseng, Hung-Yi Lee

Current state-of-the-art (SOTA) codec-based audio synthesis systems can mimic anyone's voice with just a 3-second sample from that specific unseen speaker. Unfortunately, malicious attackers may exploit these technologie…

Audio SynthesisFace SwappingSpeech Synthesis

Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion

2025-06-02 · Ajinkya Kulkarni, Sandipana Dowerah, Tanel Alumae, Mathew Magimai. -Doss

Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work prop…

Face SwappingMetric Learning