paper-with-me

Papers

CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection

2024-06-04 · Yongyi Zang, Jiatong Shi, You Zhang, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Shengyuan Xu, Wenxiao Zhao, Jing Guo, Tomoki Toda, Zhiyao Duan

Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models. Current SVDD datasets face challenges due to limited controllability, diversity in deepfake methods, and licensing restrictions. Addressing these gaps, we introduce CtrSVDD, a large-scale, diverse collection of bonafide and deepfake singing vocals. These vocals are synthesized using state-of-the-art methods from publicly accessible singing voice datasets. CtrSVDD includes 47.64 hours of bonafide and 260.34 hours of deepfake singing vocals, spanning 14 deepfake methods and involving 164 singer identities. We also present a baseline system with flexible front-end features, evaluated against a structured train/dev/eval split. The experiments show the importance of feature selection and highlight a need for generalization towards deepfake methods that deviate further from training distribution. The CtrSVDD dataset and baselines are publicly accessible.

📄 PDF Abstract BibTeX arXiv:2406.02438

Code (2)

svddchallenge/ctrsvdd2024_baseline 공식 구현 pytorch
yaselley/ssl_layerwise_deepfake pytorch

Tasks

DeepFake DetectionDiversityFace Swappingfeature selectionSinging Voice Synthesis

Methods 이 논문이 사용한 방법론

Feature Selection Feature selection, also known as variable selection, attribute selection or variable subset selection, is the process of selecting a subset of relevant features (variables,…

Similar Papers 제목 키워드 기반

SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge

2024-08-28 · You Zhang, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto 외

With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to advance research in identifying AI-gene…

DeepFake DetectionFace SwappingSinging Voice Synthesis

Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024

2024-09-03 · Anmol Guragain, Tianchi Liu, Zihan Pan, Hardik B. Sailor 외

This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generat…

DeepFake DetectionFace SwappingVoice Anti-spoofing

Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-spoofing

2025-04-08 · Tianchi Liu, Duc-Tuan Truong, Rohan Kumar Das, Kong Aik Lee 외

Speech foundation models have significantly advanced various speech-related tasks by providing exceptional representation capabilities. However, their high-dimensional output features often create a mismatch with downstr…

DeepFake DetectionDimensionality ReductionFace Swapping

iTRIALSPACE: Programmable Virtual Lesion Trials for Controlled Evaluation of Lung CT Models

2026-05-07 · Fakrul Islam Tushar, Umme Hafsa Momy, Joseph Y. Lo, Geoffrey D. Rubin arxiv

We introduce iTRIALSPACE, a programmable evaluation framework for controlled assessment of lung CT models. Standard benchmarks are static retrospective collections that entangle lesion size, lobe prevalence, anatomy, and…

StereoGenBench: A Synthetic Multi-Camera Benchmark for Stereo Generation under Controlled Baseline Regimes

2026-05-22 · Yangzhi Cui, Feng Qiao, Nathan Jacobs arxiv

Stereo image and video generation, stereo geometry estimation, and condition-controlled view synthesis require paired data in which the variables that determine binocular geometry -- camera baseline, intrinsics, scene de…

Video Generation