Efficient Personalized Speech Enhancement through Self-Supervised Learning
This work presents self-supervised learning methods for developing monaural speaker-specific (i.e., personalized) speech enhancement models. While generalist models must broadly address many speakers, specialist models can adapt their enhancement function towards a particular speaker's voice, expecting to solve a narrower problem. Hence, specialists are capable of achieving more optimal performance in addition to reducing computational complexity. However, naive personalization methods can require clean speech from the target user, which is inconvenient to acquire, e.g., due to subpar recording conditions. To this end, we pose personalization as either a zero-shot task, in which no additional clean speech of the target speaker is used for training, or a few-shot learning task, in which the goal is to minimize the duration of the clean speech used for transfer learning. With this paper, we propose self-supervised learning methods as a solution to both zero- and few-shot personalization tasks. The proposed methods are designed to learn the personalized speech features from unlabeled data (i.e., in-the-wild noisy recordings from the target user) without knowing the corresponding clean sources. Our experiments investigate three different self-supervised learning mechanisms. The results show that self-supervised models achieve zero-shot and few-shot personalization using fewer model parameters and less clean data from the target user, achieving the data efficiency and model compression goals.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot LearningModel CompressionSelf-Supervised LearningSpeech EnhancementTransfer LearningSimilar Papers 제목 키워드 기반
Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification
Training personalized speech enhancement models is innately a no-shot learning problem due to privacy constraints and limited access to noise-free speech from the target user. If there is an abundance of unlabeled noisy …
Data AugmentationDenoisingPrivacy PreservingSelf-Supervised Learning+1Self-Supervised Learning from Contrastive Mixtures for Personalized Speech Enhancement
This work explores how self-supervised learning can be universally used to discover speaker-specific features towards enabling personalized speech enhancement models. We specifically address the few-shot learning scenari…
Contrastive LearningFew-Shot LearningModel CompressionSelf-Supervised Learning+1BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportu…
DenoisingSelf-Supervised LearningSpeech EnhancementPerceive and predict: self-supervised speech representation based loss functions for speech enhancement
Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the d…
Speech EnhancementA Framework for Unified Real-time Personalized and Non-Personalized Speech Enhancement
In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning i…
Multi-Task LearningSpeech Enhancement