paper-with-me

Papers

AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection

2026-03-27 · Hai-Son Nguyen-Le, Hung-Cuong Nguyen-Thanh, Nhien-An Le-Khac, Dinh-Thuc Nguyen, Hong-Hanh Nguyen-Le arxiv

The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45\%, including a significant reduction to 1.23\% on WaveFake and 2.70\% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code is publicly available at https://github.com/NguyenLeHaiSonGit/AFSS.

📄 PDF Abstract BibTeX arXiv:2603.26856

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake Detection

Similar Papers 제목 키워드 기반

NAFSSR: Stereo Image Super-Resolution Using NAFNet

2022-04-19 · Xiaojie Chu, Liangyu Chen, Wenqing Yu

Stereo image super-resolution aims at enhancing the quality of super-resolution results by utilizing the complementary information provided by binocular systems. To obtain reasonable performance, most methods focus on fi…

Image RestorationImage Super-ResolutionStereo Image Super-ResolutionSuper-Resolution

DELINE8K: A Synthetic Data Pipeline for the Semantic Segmentation of Historical Documents

2024-04-30 · Taylor Archibald, Tony Martinez

Document semantic segmentation is a promising avenue that can facilitate document analysis tasks, including optical character recognition (OCR), form classification, and document editing. Although several synthetic datas…

8kDiversityFormOptical Character Recognition+3

TAFSSL: Task-Adaptive Feature Sub-Space Learning for few-shot classification

2020-03-14 · ECCV 2020 8 · Moshe Lichtenstein, Prasanna Sattigeri, Rogerio Feris, Raja Giryes 외

The field of Few-Shot Learning (FSL), or learning from very few (typically $1$ or $5$) examples per novel class (unseen during training), has received a lot of attention and significant performance advances in the recent…

DiversityFew-Shot LearningGeneral Classification

SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps

2022-01-13 · Carlos Rodriguez-Pardo, Elena Garces

We present SeamlessGAN, a method capable of automatically generating tileable texture maps from a single input exemplar. In contrast to most existing methods, focused solely on solving the synthesis problem, our work tac…

Image GenerationTexture Synthesisvalid

A Post Auto-regressive GAN Vocoder Focused on Spectrum Fracture

2022-04-12 · Zhenxing Lu, Mengnan He, Ruixiong Zhang, Caixia Gong

Generative adversarial networks (GANs) have been indicated their superiority in usage of the real-time speech synthesis. Nevertheless, most of them make use of deep convolutional layers as their backbone, which may cause…

Speech Synthesis