paper-with-me

홈 › Papers

Personalized Speech Enhancement Without a Separate Speaker Embedding Model

2024-06-14 · Tanel Pärnamaa, Ando Saabas

Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to extract a vector representation of the speaker from enrollment audio, which adds complexity to the training and deployment process. We propose to use the internal representation of the PSE model itself as the speaker embedding, thereby avoiding the need for a separate model. We show that our approach performs equally well or better than the standard method of using a pre-trained speaker embedding model on noise suppression and echo cancellation tasks. Moreover, our approach surpasses the ICASSP 2023 Deep Noise Suppression Challenge winner by 0.15 in Mean Opinion Score.

📄 PDF Abstract BibTeX arXiv:2406.09928

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Personalized PercepNet: Real-time, Low-complexity Target Voice Separation and Enhancement

2021-06-08 · Ritwik Giri, Shrikant Venkataramani, Jean-Marc Valin, Umut Isik 외

The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present…

A Framework for Unified Real-time Personalized and Non-Personalized Speech Enhancement

2023-02-23 · Zhepei Wang, Ritwik Giri, Devansh Shah, Jean-Marc Valin 외

In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning i…

Multi-Task LearningSpeech Enhancement

Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification

2021-04-05 · Aswin Sivaraman, Sunwoo Kim, Minje Kim

Training personalized speech enhancement models is innately a no-shot learning problem due to privacy constraints and limited access to noise-free speech from the target user. If there is an abundance of unlabeled noisy …

Data AugmentationDenoisingPrivacy PreservingSelf-Supervised Learning+1

Personalized speech enhancement combining band-split RNN and speaker attentive module

2023-02-20 · Xiaohuai Le, Li Chen, Chao He, Yiqing Guo 외

Target speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of conca…

Speech Enhancement

Efficient Personalized Speech Enhancement through Self-Supervised Learning

2021-04-05 · Aswin Sivaraman, Minje Kim

This work presents self-supervised learning methods for developing monaural speaker-specific (i.e., personalized) speech enhancement models. While generalist models must broadly address many speakers, specialist models c…

Few-Shot LearningModel CompressionSelf-Supervised LearningSpeech Enhancement+1