Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection
Supervised contrastive learning (SupCon) is widely used to shape representations, but has seen limited targeted study for audio deepfake detection. Existing work typically combines contrastive terms with broader pipelines; however, the focus on SupCon itself is missing. In this work, we run a controlled study on wav2vec2 XLS-R (300M) that varies (i) similarity in SupCon (cosine vs angular similarity derived from the hyperspherical angle) and (ii) negative scaling using a warm-started global cross-batch queue. Stage 1 fine-tunes the encoder and projection head with SupCon; Stage 2 freezes them and trains a linear classifier with BCE. Trained on ASVspoof 2019 LA and evaluated on ASV19 eval plus ITW and ASVspoof 2021 DF/LA, Cosine SupCon with a delayed queue achieves the best ITW EER (8.29%) and pooled EER (4.44), while angular similarity performs strongly without queued negatives (ITW 8.70), indicating reduced reliance on large negative sets.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio Deepfake DetectionContrastive LearningSimilar Papers 제목 키워드 기반
Unsupervised hard Negative Augmentation for contrastive learning
We present Unsupervised hard Negative Augmentation (UNA), a method that generates synthetic negative instances based on the term frequency-inverse document frequency (TF-IDF) retrieval model. UNA uses TF-IDF scores to as…
Contrastive LearningData AugmentationRetrievalSemantic Textual Similarity+3Investigating the Role of Negatives in Contrastive Representation Learning
Noise contrastive learning is a popular technique for unsupervised representation learning. In this approach, a representation is obtained via reduction to supervised learning, where given a notion of semantic similarity…
Contrastive LearningData AugmentationRepresentation LearningSemantic Similarity+1Design of the topology for contrastive visual-textual alignment
Cosine similarity is the common choice for measuring the distance between the feature representations in contrastive visual-textual alignment learning. However, empirically a learnable softmax temperature parameter is re…
Contrastive LearningImage-to-Text RetrievalRetrievalText Retrieval+2Dynamically Scaled Temperature in Self-Supervised Contrastive Learning
In contemporary self-supervised contrastive algorithms like SimCLR, MoCo, etc., the task of balancing attraction between two semantically similar samples and repulsion between two samples of different classes is primaril…
Contrastive LearningSelf-Supervised LearningSNCSE: Contrastive Learning for Unsupervised Sentence Embedding with Soft Negative Samples
Unsupervised sentence embedding aims to obtain the most appropriate embedding for a sentence to reflect its semantic. Contrastive learning has been attracting developing attention. For a sentence, current models utilize …
Contrastive LearningData AugmentationNegationSemantic Similarity+5