paper-with-me

Papers

DiffSSD: A Diffusion-Based Dataset For Speech Forensics

2024-09-19 · Kratika Bhagtani, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp

Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed. Many of these detectors are trained on datasets which do not include diffusion-based synthesizers. In this paper, we demonstrate that existing detectors trained on one such dataset, ASVspoof2019, do not perform well in detecting synthetic speech from recent diffusion-based synthesizers. We propose the Diffusion-Based Synthetic Speech Dataset (DiffSSD), a dataset consisting of about 200 hours of labeled speech, including synthetic speech generated by 8 diffusion-based open-source and 2 commercial generators. We also examine the performance of existing synthetic speech detectors on DiffSSD in both closed-set and open-set scenarios. The results highlight the importance of this dataset in detecting synthetic speech generated from recent open-source and commercial speech generators.

📄 PDF Abstract BibTeX arXiv:2409.13049

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HISPASpoof: A New Dataset For Spanish Speech Forensics

2025-09-11 · Maria Risques, Kratika Bhagtani, Amit Kumar Singh Yadav, Edward J. Delp arxiv

Zero-shot Voice Cloning (VC) and Text-to-Speech (TTS) methods have advanced rapidly, enabling the generation of highly realistic synthetic speech and raising serious concerns about their misuse. While numerous detectors …

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

2025-08-13 · Yachao Liang, Min Yu, Gang Li, Jianguo Jiang 외 arxiv

Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leverag…

Representation Learning

Robust Spoofed Speech Detection via Temporal Pyramid Modeling

2026-06-15 · Mahtab Masoudi Nezhad, Nima Karimian arxiv

Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset generalization remaining a major limitation. This work we propose a Temporal Pyramid Ad…

Voice Conversion

Diffusion models meet image counter-forensics

2023-11-22 · Matías Tailanian, Marina Gardella, Álvaro Pardo, Pablo Musé

From its acquisition in the camera sensors to its storage, different operations are performed to generate the final image. This pipeline imprints specific traces into the image to form a natural watermark. Tampering with…

Adversarial Purification

Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis

2024-12-12 · Zhoulin Ji, Chenhao Lin, Hang Wang, Chao Shen

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on sp…

Misinformation