paper-with-me

Papers

Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering

2023-05-18 · Heng-Jui Chang, Alexander H. Liu, James Glass

Self-supervised speech representation models have succeeded in various tasks, but improving them for content-related problems using unlabeled data is challenging. We propose speaker-invariant clustering (Spin), a novel self-supervised learning method that clusters speech representations and performs swapped prediction between the original and speaker-perturbed utterances. Spin disentangles speaker information and preserves content representations with just 45 minutes of fine-tuning on a single GPU. Spin improves pre-trained networks and outperforms prior methods in speech recognition and acoustic unit discovery.

📄 PDF Abstract BibTeX arXiv:2305.11072

Code (1)

vectominist/spin 공식 구현 pytorch

Tasks

Acoustic Unit DiscoveryClusteringGPUSelf-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations

2024-03-10 · Amit Meghanani, Thomas Hain

There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used…

Automatic Speech RecognitionData AugmentationGPUPhoneme Recognition+3

Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations

2026-01-28 · Amit Meghanani, Thomas Hain arxiv

Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representati…

Self-Supervised LearningSpeech Enhancement

LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks

2024-06-13 · Amit Meghanani, Thomas Hain

Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-r…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)GPUPhoneme Recognition+3

Injecting Text and Cross-lingual Supervision in Few-shot Learning from Self-Supervised Models

2021-10-10 · Matthew Wiesner, Desh Raj, Sanjeev Khudanpur

Self-supervised model pre-training has recently garnered significant interest, but relatively few efforts have explored using additional resources in fine-tuning these models. We demonstrate how universal phoneset acoust…

Few-Shot Learning

Layer-wise Analysis of a Self-supervised Speech Representation Model

2021-07-10 · Ankita Pasad, Ju-chieh Chou, Karen Livescu

Recently proposed self-supervised learning approaches have been successful for pre-training speech representation models. The utility of these learned representations has been observed empirically, but not much has been …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1