Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
Recent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts. This improvement could be due to the powerful modeling ability of the multi-head self-attention (MHSA) in the Transformer model, which learns the temporal relationship of each input token. However, artifacts of synthetic speech can be located in specific regions of both frequency channels and temporal segments, while MHSA neglects this temporal-channel dependency of the input sequence. In this work, we proposed a Temporal-Channel Modeling (TCM) module to enhance MHSA's capability for capturing temporal-channel dependencies. Experimental results on the ASVspoof 2021 show that with only 0.03M additional parameters, the TCM module can outperform the state-of-the-art system by 9.25% in EER. Further ablation study reveals that utilizing both temporal and channel information yields the most improvement for detecting synthetic speech.
Code (1)
Tasks
Audio Deepfake DetectionSynthetic Speech DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
GTA: Global Temporal Attention for Video Action Understanding
Self-attention learns pairwise interactions to model long-range dependencies, yielding great improvements for video action recognition. In this paper, we seek a deeper understanding of self-attention for temporal modelin…
Action RecognitionAction UnderstandingTemporal Action LocalizationRECTOR: Masked Region-Channel-Temporal Modeling for Affective and Cognitive Representation Learning
Affective and cognitive disorders manifest as distributed, time-varying brain network dynamics across regions, channels, and time, challenging robust representation learning from EEG/sEEG for clinical diagnosis. We propo…
EEG Emotion RecognitionRepresentation LearningHierarchical Separable Video Transformer for Snapshot Compressive Imaging
Transformers have achieved the state-of-the-art performance on solving the inverse problem of Snapshot Compressive Imaging (SCI) for video, whose ill-posedness is rooted in the mixed degradation of spatial masking and te…
Inductive BiasLong-range modelingLightweight Temporal Self-Attention for Classifying Satellite Image Time Series
The increasing accessibility and precision of Earth observation satellite data offers considerable opportunities for industrial and state actors alike. This calls however for efficient methods able to process time-series…
Earth ObservationTime SeriesTime Series AnalysisTime Series ClassificationMILAAP: Mobile Link Allocation via Attention-based Prediction
Channel hopping (CS) communication systems must adapt to interference changes in the wireless network and to node mobility for maintaining throughput efficiency. Optimal scheduling requires up-to-date network state infor…
PredictionScheduling