paper-with-me

홈 › Papers

Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models

2024-09-21 · Orchid Chetia Phukan, Sarthak Jain, Swarup Ranjan Behera, Arun Balaji Buduru, Rajesh Sharma, S. R Mahadeva Prasanna

In this study, for the first time, we extensively investigate whether music foundation models (MFMs) or speech foundation models (SFMs) work better for singing voice deepfake detection (SVDD), which has recently attracted attention in the research community. For this, we perform a comprehensive comparative study of state-of-the-art (SOTA) MFMs (MERT variants and music2vec) and SFMs (pre-trained for general speech representation learning as well as speaker recognition). We show that speaker recognition SFM representations perform the best amongst all the foundation models (FMs), and this performance can be attributed to its higher efficacy in capturing the pitch, tone, intensity, etc, characteristics present in singing voices. To our end, we also explore the fusion of FMs for exploiting their complementary behavior for improved SVDD, and we propose a novel framework, FIONA for the same. With FIONA, through the synchronization of x-vector (speaker recognition SFM) and MERT-v1-330M (MFM), we report the best performance with the lowest Equal Error Rate (EER) of 13.74 %, beating all the individual FMs as well as baseline FM fusions and achieving SOTA results.

📄 PDF Abstract BibTeX arXiv:2409.14131

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake DetectionFace SwappingRepresentation LearningSpeaker RecognitionSpeech Representation Learning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

2025-06-03 · Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera 외

In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they …

Face Swapping

SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan

2024-05-08 · You Zhang, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto 외

The rapid advancement of AI-generated singing voices, which now closely mimic natural human singing and align seamlessly with musical scores, has led to heightened concerns for artists and the music industry. Unlike spok…

DeepFake DetectionFace Swapping

SingFake: Singing Voice Deepfake Detection

2023-09-14 · Yongyi Zang, You Zhang, Mojtaba Heydari, Zhiyao Duan

The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage. Unlike synthesized speech, synthesized singing voices are typically released in songs c…

DeepFake DetectionFace SwappingSinging Voice SynthesisSynthetic Speech Detection

Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024

2024-09-03 · Anmol Guragain, Tianchi Liu, Zihan Pan, Hardik B. Sailor 외

This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generat…

DeepFake DetectionFace SwappingVoice Anti-spoofing

Singing Voice Graph Modeling for SingFake Detection

2024-06-05 · Xuanjun Chen, Haibin Wu, Jyh-Shing Roger Jang, Hung-Yi Lee

Detecting singing voice deepfakes, or SingFake, involves determining the authenticity and copyright of a singing voice. Existing models for speech deepfake detection have struggled to adapt to unseen attacks in this uniq…

DeepFake DetectionFace SwappingRhythm