Synthetic Speech Detection Based on Temporal Consistency and Distribution of Speaker Features
Current synthetic speech detection (SSD) methods perform well on certain datasets but still face issues of robustness and interpretability. A possible reason is that these methods do not analyze the deficiencies of synthetic speech. In this paper, the flaws of the speaker features inherent in the text-to-speech (TTS) process are analyzed. Differences in the temporal consistency of intra-utterance speaker features arise due to the lack of fine-grained control over speaker features in TTS. Since the speaker representations in TTS are based on speaker embeddings extracted by encoders, the distribution of inter-utterance speaker features differs between synthetic and bonafide speech. Based on these analyzes, an SSD method based on temporal consistency and distribution of speaker features is proposed. On one hand, modeling the temporal consistency of intra-utterance speaker features can aid speech anti-spoofing. On the other hand, distribution differences in inter-utterance speaker features can be utilized for SSD. The proposed method offers low computational complexity and performs well in both cross-dataset and silence trimming scenarios.
Code (0)
등록된 구현이 없습니다.
Tasks
Synthetic Speech Detectiontext-to-speechText to SpeechMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what extent can low error rates on academic b…
BenchmarkingSelf-Supervised LearningSynthetic Speech Detectiontext-to-speech+1Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
Recent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts. This improvement could be due to the powerful modeling ability of th…
Audio Deepfake DetectionSynthetic Speech DetectionCompression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods …
MisinformationSynthetic Speech DetectionSpeech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on sp…
MisinformationDelving into the Frequency: Temporally Consistent Human Motion Transfer in the Fourier Space
Human motion transfer refers to synthesizing photo-realistic and temporally coherent videos that enable one person to imitate the motion of others. However, current synthetic videos suffer from the temporal inconsistency…
DeepFake DetectionFace Swapping