paper-with-me

홈 › Papers

Test-Time Adaptation Toward Personalized Speech Enhancement: Zero-Shot Learning with Knowledge Distillation

2021-05-08 · Sunwoo Kim, Minje Kim

In realistic speech enhancement settings for end-user devices, we often encounter only a few speakers and noise types that tend to reoccur in the specific acoustic environment. We propose a novel personalized speech enhancement method to adapt a compact denoising model to the test-time specificity. Our goal in this test-time adaptation is to utilize no clean speech target of the test speaker, thus fulfilling the requirement for zero-shot learning. To complement the lack of clean utterance, we employ the knowledge distillation framework. Instead of the missing clean utterance target, we distill the more advanced denoising results from an overly large teacher model, and use it as the pseudo target to train the small student model. This zero-shot learning procedure circumvents the process of collecting users' clean speech, a process that users are reluctant to comply due to privacy concerns and technical difficulty of recording clean voice. Experiments on various test-time conditions show that the proposed personalization method achieves significant performance gains compared to larger baseline networks trained from a large speaker- and noise-agnostic datasets. In addition, since the compact personalized models can outperform larger general-purpose models, we claim that the proposed method performs model compression with no loss of denoising performance.

📄 PDF Abstract BibTeX arXiv:2105.03544

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingKnowledge DistillationModel CompressionSpecificitySpeech EnhancementTest-time AdaptationZero-Shot Learning

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Zero-Shot Personalized Speech Enhancement through Speaker-Informed Model Selection

2021-05-08 · Aswin Sivaraman, Minje Kim

This paper presents a novel zero-shot learning approach towards personalized speech enhancement through the use of a sparsely active ensemble model. Optimizing speech denoising systems towards a particular test-time spea…

ClusteringDenoisingModel SelectionSpeech Denoising+2

Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification

2021-04-05 · Aswin Sivaraman, Sunwoo Kim, Minje Kim

Training personalized speech enhancement models is innately a no-shot learning problem due to privacy constraints and limited access to noise-free speech from the target user. If there is an abundance of unlabeled noisy …

Data AugmentationDenoisingPrivacy PreservingSelf-Supervised Learning+1

A Framework for Unified Real-time Personalized and Non-Personalized Speech Enhancement

2023-02-23 · Zhepei Wang, Ritwik Giri, Devansh Shah, Jean-Marc Valin 외

In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning i…

Multi-Task LearningSpeech Enhancement

The Potential of Neural Speech Synthesis-based Data Augmentation for Personalized Speech Enhancement

2022-11-14 · Anastasia Kuznetsova, Aswin Sivaraman, Minje Kim

With the advances in deep learning, speech enhancement systems benefited from large neural network architectures and achieved state-of-the-art quality. However, speaker-agnostic methods are not always desirable, both in …

Data AugmentationSpeech EnhancementSpeech Synthesis

Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention

2020-02-14 · Yuma Koizumi, Kohei Yatabe, Marc Delcroix, Yoshiki Masuyama 외

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studi…

Multi-Task LearningSpeaker IdentificationSpeech Enhancementspeech-recognition+1