paper-with-me

Papers

Diffuse or Confuse: A Diffusion Deepfake Speech Dataset

2024-10-09 · Anton Firc, Kamil Malinka, Petr Hanáček

Advancements in artificial intelligence and machine learning have significantly improved synthetic speech generation. This paper explores diffusion models, a novel method for creating realistic synthetic speech. We create a diffusion dataset using available tools and pretrained models. Additionally, this study assesses the quality of diffusion-generated deepfakes versus non-diffusion ones and their potential threat to current deepfake detection systems. Findings indicate that the detection of diffusion-based deepfakes is generally comparable to non-diffusion deepfakes, with some variability based on detector architecture. Re-vocoding with diffusion vocoders shows minimal impact, and the overall speech quality is comparable to non-diffusion methods.

📄 PDF Abstract BibTeX arXiv:2410.06796

Code (1)

AntonFirc/diffusion-deepfake-speech-dataset 공식 구현

Tasks

DeepFake DetectionFace Swapping

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models

2025-05-16 · Danilo de Oliveira, Julius Richter, Tal Peer, Timo Gerkmann

We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

The DiffuseStyleGesture+ entry to the GENEA Challenge 2023

2023-08-26 · Sicheng Yang, Haiwei Xue, Zhensong Zhang, Minglei Li 외

In this paper, we introduce the DiffuseStyleGesture+, our solution for the Generation and Evaluation of Non-verbal Behavior for Embodied Agents (GENEA) Challenge 2023, which aims to foster the development of realistic, a…

A Study on Speech Enhancement Based on Diffusion Probabilistic Model

2021-07-25 · Yen-Ju Lu, Yu Tsao, Shinji Watanabe

Diffusion probabilistic models have demonstrated an outstanding capability to model natural images and raw audio waveforms through a paired diffusion and reverse processes. The unique property of the reverse process (nam…

Speech Enhancement

FaceDiffuser: Speech-Driven 3D Facial Animation Synthesis Using Diffusion

2023-09-20 · Stefan Stan, Kazi Injamamul Haque, Zerrin Yumak

Speech-driven 3D facial animation synthesis has been a challenging task both in industry and research. Recent methods mostly focus on deterministic deep learning methods meaning that given a speech input, the output is a…

3D Face Animation

DeePhy: On Deepfake Phylogeny

2022-09-19 · Kartik Narayan, Harsh Agarwal, Kartik Thakral, Surbhi Mittal 외

Deepfake refers to tailored and synthetically generated videos which are now prevalent and spreading on a large scale, threatening the trustworthiness of the information available online. While existing datasets contain …

DeepFake DetectionFace Swapping