paper-with-me

홈 › Papers

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

2025-06-03 · Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Priyabrata Mallick, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma

In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they are better equipped for capturing subtle source-specific characteristics-such as unique timbre, pitch manipulation, or synthesis artifacts of each singing voice deepfake source due to their cross-modality pre-training. Our experiments with MMFMs, speech foundation models and music foundation models verify the hypothesis that MMFMs are the most effective for SVDSA. Furthermore, inspired from related research, we also explore fusion of foundation models (FMs) for improved SVDSA. To this end, we propose a novel framework, COFFE which employs Chernoff Distance as novel loss function for effective fusion of FMs. Through COFFE with the symphony of MMFMs, we attain the topmost performance in comparison to all the individual FMs and baseline fusion methods.

📄 PDF Abstract BibTeX arXiv:2506.03364

Code (1)

Helix-IIIT-Delhi/COFFE-Singing_Voice_Deepfake 공식 구현

Tasks

Face Swapping

Similar Papers 제목 키워드 기반

SingFake: Singing Voice Deepfake Detection

2023-09-14 · Yongyi Zang, You Zhang, Mojtaba Heydari, Zhiyao Duan

The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage. Unlike synthesized speech, synthesized singing voices are typically released in songs c…

DeepFake DetectionFace SwappingSinging Voice SynthesisSynthetic Speech Detection

SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge

2024-08-28 · You Zhang, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto 외

With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to advance research in identifying AI-gene…

DeepFake DetectionFace SwappingSinging Voice Synthesis

CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection

2024-06-04 · Yongyi Zang, Jiatong Shi, You Zhang, Ryuichi Yamamoto 외

Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models. Current SVDD datasets face challenges due to limited controllability, diversity in deepfake me…

DeepFake DetectionDiversityFace Swappingfeature selection+1

Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024

2024-09-03 · Anmol Guragain, Tianchi Liu, Zihan Pan, Hardik B. Sailor 외

This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generat…

DeepFake DetectionFace SwappingVoice Anti-spoofing

SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan

2024-05-08 · You Zhang, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto 외

The rapid advancement of AI-generated singing voices, which now closely mimic natural human singing and align seamlessly with musical scores, has led to heightened concerns for artists and the music industry. Unlike spok…

DeepFake DetectionFace Swapping