paper-with-me

Papers

Speakerfilter-Pro: an improved target speaker extractor combines the time domain and frequency domain

2020-10-25 · Shulin He, Hao Li, Xueliang Zhang

This paper introduces an improved target speaker extractor, referred to as Speakerfilter-Pro, based on our previous Speakerfilter model. The Speakerfilter uses a bi-direction gated recurrent unit (BGRU) module to characterize the target speaker from anchor speech and use a convolutional recurrent network (CRN) module to separate the target speech from a noisy signal.Different from the Speakerfilter, the Speakerfilter-Pro sticks a WaveUNet module in the beginning and the ending, respectively. The WaveUNet has been proven to have a better ability to perform speech separation in the time domain. In order to extract the target speaker information better, the complex spectrum instead of the magnitude spectrum is utilized as the input feature for the CRN module. Experiments are conducted on the two-speaker dataset (WSJ0-mix2) which is widely used for speaker extraction. The systematic evaluation shows that the Speakerfilter-Pro outperforms the Speakerfilter and other baselines, and achieves a signal-to-distortion ratio (SDR) of 14.95 dB.

📄 PDF Abstract BibTeX arXiv:2010.13053

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Methods 이 논문이 사용한 방법론

CRN Conditional Relation Network, or CRN, is a building block to construct more sophisticated structures for representation and reasoning over video. CRN takes as input an…

Similar Papers 제목 키워드 기반

Mitigating Non-Target Speaker Bias in Guided Speaker Embedding

2025-06-14 · Shota Horiguchi, Takanori Ashihara, Marc Delcroix, Atsushi Ando 외

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target sp…

Speaker Verification

High-resolution embedding extractor for speaker diarisation

2022-11-08 · Hee-Soo Heo, Youngki Kwon, Bong-Jin Lee, You Jin Kim 외

Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the slid…

Vocal Bursts Intensity Prediction

The Royalflush System for VoxCeleb Speaker Recognition Challenge 2022

2022-09-19 · Jingguang Tian, Xinhui Hu, Xinkang Xu

In this technical report, we describe the Royalflush submissions for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). Our submissions contain track 1, which is for supervised speaker verification and track 3,…

ClusteringDomain AdaptationSpeaker RecognitionSpeaker Verification

A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures

2023-06-01 · Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă, Rama Doddipatla 외

We introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture. To allow for supervised training, a teacher-student approach is employed: the teacher c…

Challenging margin-based speaker embedding extractors by using the variational information bottleneck

2024-06-18 · Themos Stafylakis, Anna Silnova, Johan Rohdin, Oldrich Plchot 외

Speaker embedding extractors are typically trained using a classification loss over the training speakers. During the last few years, the standard softmax/cross-entropy loss has been replaced by the margin-based losses, …

Speaker Recognition