Simultaneous Speech Extraction for Multiple Target Speakers under the Meeting Scenarios
The common target speech separation directly estimate the target source, ignoring the interrelationship between different speakers at each frame. We propose a multiple-target speech separation model (MTSS) to simultaneously extract each speaker's voice from the mixed speech rather than just optimally estimating the target source. Moreover, we propose a speaker diarization (SD) aware MTSS system (SD-MTSS), which consists of a SD module and MTSS module. By exploiting the TSVAD decision and the estimated mask, our SD-MTSS model can extract the speech signal of each speaker concurrently in a conversational recording without additional enrollment audio in advance. Experimental results show that our MTSS model achieves 1.38dB SDR, 1.34dB SI-SDR, and 0.13 PESQ improvements over the baseline on the WSJ0-2mix-extr dataset, respectively. The SD-MTSS system makes 19.2% relative speaker dependent character error rate (CER) reduction on the Alimeeting dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity Detectionspeaker-diarizationSpeaker DiarizationSpeech ExtractionSpeech SeparationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Coarse-to-Fine Recursive Speech Separation for Unknown Number of Speakers
The vast majority of speech separation methods assume that the number of speakers is known in advance, hence they are specific to the number of speakers. By contrast, a more realistic and challenging task is to separate …
Speech SeparationTarget Speaker ExtractionImproving speaker discrimination of target speech extraction with time-domain SpeakerBeam
Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation u…
Speaker IdentificationSpeech ExtractionSpeech SeparationListen to Extract: Onset-Prompted Target Speaker Extraction
We propose $\textit{listen to extract}$ (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting…
Target Speaker ExtractionTarget Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean au…
Target Speaker ExtractionSpatially Selective Deep Non-linear Filters for Speaker Extraction
In a scenario with multiple persons talking simultaneously, the spatial characteristics of the signals are the most distinct feature for extracting the target signal. In this work, we develop a deep joint spatial-spectra…
Speech Separation