Feature Normalization for Fine-tuning Self-Supervised Models in Speech Enhancement
Large, pre-trained representation models trained using self-supervised learning have gained popularity in various fields of machine learning because they are able to extract high-quality salient features from input data. As such, they have been frequently used as base networks for various pattern classification tasks such as speech recognition. However, not much research has been conducted on applying these types of models to the field of speech signal generation. In this paper, we investigate the feasibility of using pre-trained speech representation models for a downstream speech enhancement task. To alleviate mismatches between the input features of the pre-trained model and the target enhancement model, we adopt a novel feature normalization technique to smoothly link these modules together. Our proposed method enables significant improvements in speech quality compared to baselines when combined with various types of pre-trained speech models.
Code (0)
등록된 구현이 없습니다.
Tasks
Self-Supervised LearningSpeech Enhancementspeech-recognitionSpeech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Improved transferability of self-supervised learning models through batch normalization finetuning
Abundance of unlabelled data and advances in Self-Supervised Learning (SSL) have made it the preferred choice in many transfer learning scenarios. Due to the rapid and ongoing development of SSL approaches, practitioners…
ClassificationFew-Shot LearningSelf-Supervised LearningTransfer LearningEvaluating the fairness of fine-tuning strategies in self-supervised learning
In this work we examine how fine-tuning impacts the fairness of contrastive Self-Supervised Learning (SSL) models. Our findings indicate that Batch Normalization (BN) statistics play a crucial role, and that updating onl…
FairnessSelf-Supervised LearningObjectives Matter: Understanding the Impact of Self-Supervised Objectives on Vision Transformer Representations
Joint-embedding based learning (e.g., SimCLR, MoCo, DINO) and reconstruction-based learning (e.g., BEiT, SimMIM, MAE) are the two leading paradigms for self-supervised learning of vision transformers, but they differ sub…
Self-Supervised LearningSpecificityAC-Norm: Effective Tuning for Medical Image Analysis via Affine Collaborative Normalization
Driven by the latest trend towards self-supervised learning (SSL), the paradigm of "pretraining-then-finetuning" has been extensively explored to enhance the performance of clinical applications with limited annotations.…
Cardiac SegmentationLung Nodule SegmentationMedical Image AnalysisRetinal Vessel Segmentation+3RINO: Renormalization Group Invariance with No Labels
A common challenge with supervised machine learning (ML) in high energy physics (HEP) is the reliance on simulations for labeled data, which can often mismodel the underlying collision or detector response. To help mitig…
Self-Supervised Learning