Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect synthetic speech have been proposed. Existing methods often overfit on one dataset and their performance reduces substantially in practical scenarios such as detecting synthetic speech shared on social platforms. In this paper we propose, Patched Spectrogram Synthetic Speech Detection Transformer (PS3DT), a synthetic speech detector that converts a time domain speech signal to a mel-spectrogram and processes it in patches using a transformer neural network. We evaluate the detection performance of PS3DT on ASVspoof2019 dataset. Our experiments show that PS3DT performs well on ASVspoof2019 dataset compared to other approaches using spectrogram for synthetic speech detection. We also investigate generalization performance of PS3DT on In-the-Wild dataset. PS3DT generalizes well than several existing methods on detecting synthetic speech from an out-of-distribution dataset. We also evaluate robustness of PS3DT to detect telephone quality synthetic speech and synthetic speech shared on social platforms (compressed speech). PS3DT is robust to compression and can detect telephone quality synthetic speech better than several existing methods.
Code (0)
등록된 구현이 없습니다.
Tasks
MisinformationSynthetic Speech DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DSVAE: Interpretable Disentangled Representation for Synthetic Speech Detection
Tools to generate high quality synthetic speech signal that is perceptually indistinguishable from speech recorded from human speakers are easily available. Several approaches have been proposed for detecting synthetic s…
Representation LearningSynthetic Speech DetectionAudio Deepfake Detection Based on a Combination of F0 Information and Real Plus Imaginary Spectrogram Features
Recently, pioneer research works have proposed a large number of acoustic features (log power spectrogram, linear frequency cepstral coefficients, constant Q cepstral coefficients, etc.) for audio deepfake detection, obt…
Audio Deepfake DetectionDeepFake DetectionFace SwappingDiscrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition
Discrete audio representation, aka audio tokenization, has seen renewed interest driven by its potential to facilitate the application of text language modeling approaches in audio domain. To this end, various compressio…
Language ModelingLanguage ModellingQuantizationRepresentation Learning+3Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition
For many Automatic Speech Recognition (ASR) tasks audio features as spectrograms show better results than Mel-frequency Cepstral Coefficients (MFCC), but in practice they are hard to use due to a complex dimensionality o…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionDiffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found Data
Creating synthetic voices with found data is challenging, as real-world recordings often contain various types of audio degradation. One way to address this problem is to pre-enhance the speech with an enhancement model …
Speech EnhancementSpeech Synthesistext-to-speechText to Speech