Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering
Self-supervised speech representation models have succeeded in various tasks, but improving them for content-related problems using unlabeled data is challenging. We propose speaker-invariant clustering (Spin), a novel self-supervised learning method that clusters speech representations and performs swapped prediction between the original and speaker-perturbed utterances. Spin disentangles speaker information and preserves content representations with just 45 minutes of fine-tuning on a single GPU. Spin improves pre-trained networks and outperforms prior methods in speech recognition and acoustic unit discovery.
Code (1)
Tasks
Acoustic Unit DiscoveryClusteringGPUSelf-Supervised Learningspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used…
Automatic Speech RecognitionData AugmentationGPUPhoneme Recognition+3Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representati…
Self-Supervised LearningSpeech EnhancementLASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-r…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)GPUPhoneme Recognition+3Injecting Text and Cross-lingual Supervision in Few-shot Learning from Self-Supervised Models
Self-supervised model pre-training has recently garnered significant interest, but relatively few efforts have explored using additional resources in fine-tuning these models. We demonstrate how universal phoneset acoust…
Few-Shot LearningLayer-wise Analysis of a Self-supervised Speech Representation Model
Recently proposed self-supervised learning approaches have been successful for pre-training speech representation models. The utility of these learned representations has been observed empirically, but not much has been …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1