paper-with-me

Papers

Target Speech Extraction Based on Blind Source Separation and X-vector-based Speaker Selection Trained with Data Augmentation

2020-05-16 · Zhaoyi Gu, Lele Liao, Kai Chen, Jing Lu

Extracting the desired speech from a mixture is a meaningful and challenging task. The end-to-end DNN-based methods, though attractive, face the problem of generalization. In this paper, we explore a sequential approach for target speech extraction by combining blind source separation (BSS) with the x-vector based speaker recognition (SR) module. Two promising BSS methods based on source independence assumption, independent low-rank matrix analysis (ILRMA) and multi-channel variational autoencoder (MVAE), are utilized and compared. ILRMA employs nonnegative matrix factorization (NMF) to capture spectral structures of source signals and MVAE utilizes the strong modeling power of deep neural networks (DNN). However, the investigation of MVAE has been limited to the training with very few speakers and the speech signals of test speakers are usually included. We extend the training of MVAE using clean speech signals of 500 speakers to evaluate its generalization to unseen speakers. To improve the correct extraction rate, two data augmentation strategies are implemented to train the SR module. The performance of the proposed cascaded approach is investigated with test data constructed with real room impulse responses under varied environments.

📄 PDF Abstract BibTeX arXiv:2005.07976

Code (1)

annie-gu/MVAEBasedBSE 공식 구현

Tasks

blind source separationData AugmentationSpeaker RecognitionSpeech Extraction

Similar Papers 제목 키워드 기반

USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

2024-09-04 · Bang Zeng, Ming Li

Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recogniti…

Speaker RecognitionSpeech SeparationTarget Speaker Extraction

TF-MLPNet: Tiny Real-Time Neural Speech Separation

2025-08-05 · Malek Itani, Tuochao Chen, Shyamnath Gollakota arxiv

Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelera…

Speech ExtractionSpeech Separation

Analysis of impact of emotions on target speech extraction and speech separation

2022-08-15 · Ján Švec, Kateřina Žmolíková, Martin Kocour, Marc Delcroix 외

Recently, the performance of blind speech separation (BSS) and target speech extraction (TSE) has greatly progressed. Most works, however, focus on relatively well-controlled conditions using, e.g., read speech. The perf…

Speaker VerificationSpeech ExtractionSpeech Separation

Similarity-and-Independence-Aware Beamformer: Method for Target Source Extraction using Magnitude Spectrogram as Reference

2020-06-01 · Atsuo Hiroe

This study presents a novel method for source extraction, referred to as the similarity-and-independence-aware beamformer (SIBF). The SIBF extracts the target signal using a rough magnitude spectrogram as the reference s…

Speech Enhancement

Independent Vector Extraction for Fast Joint Blind Source Separation and Dereverberation

2021-02-09 · Rintaro Ikeshita, Tomohiro Nakatani

We address a blind source separation (BSS) problem in a noisy reverberant environment in which the number of microphones $M$ is greater than the number of sources of interest, and the other noise components can be approx…

blind source separation