paper-with-me

홈 › Papers

Target Confusion in End-to-end Speaker Extraction: Analysis and Approaches

2022-04-04 · Zifeng Zhao, Dongchao Yang, Rongzhi Gu, Haoran Zhang, Yuexian Zou

Recently, end-to-end speaker extraction has attracted increasing attention and shown promising results. However, its performance is often inferior to that of a blind source separation (BSS) counterpart with a similar network architecture, due to the auxiliary speaker encoder may sometimes generate ambiguous speaker embeddings. Such ambiguous guidance information may confuse the separation network and hence lead to wrong extraction results, which deteriorates the overall performance. We refer to this as the target confusion problem. In this paper, we conduct an analysis of such an issue and solve it in two stages. In the training phase, we propose to integrate metric learning methods to improve the distinguishability of embeddings produced by the speaker encoder. While for inference, a novel post-filtering strategy is designed to revise the wrong results. Specifically, we first identify these confusion samples by measuring the similarities between output estimates and enrollment utterances, after which the true target sources are recovered by a subtraction operation. Experiments show that performance improvement of more than 1dB SI-SDRi can be brought, which validates the effectiveness of our methods and emphasizes the impact of the target confusion problem.

📄 PDF Abstract BibTeX arXiv:2204.01355

Code (0)

등록된 구현이 없습니다.

Tasks

blind source separationMetric LearningSpeaker SeparationSpeech Separation

Similar Papers 제목 키워드 기반

NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection

2023-12-12 · Zexu Pan, Gordon Wichern, Francois G. Germain, Sameer Khurana 외

Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recor…

EEG

Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction

2023-12-16 · Zhaoxi Mu, Xinyu Yang, Sining Sun, Qing Yang

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic…

DisentanglementRepresentation LearningSpeech Extraction

Multi-Level Speaker Representation for Target Speaker Extraction

2024-10-21 · Ke Zhang, Junjie Li, Shuai Wang, Yangjie Wei 외

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with…

Target Speaker Extraction

X-SepFormer: End-to-end Speaker Extraction Network with Explicit Optimization on Speaker Confusion

2023-03-09 · Kai Liu, Ziqing Du, Xucheng Wan, Huan Zhou

Target speech extraction (TSE) systems are designed to extract target speech from a multi-talker mixture. The popular training objective for most prior TSE networks is to enhance reconstruction performance of extracted s…

Speech Extraction

VorTEX: Various overlap ratio for Target speech EXtraction

2026-03-16 · Ro-hoon Oh, Jihwan Seol, Bugeun Kim arxiv

Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overlapped mixtures, limiting insight into be…

Speech Extraction