Speaker activity driven neural speech extraction
Target speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have been investigated such as pre-recorded enrollment utterances, direction information, or video of the target speaker. In this paper, we explore the use of speaker activity information as an auxiliary clue for single-channel neural network-based speech extraction. We propose a speaker activity driven speech extraction neural network (ADEnet) and show that it can achieve performance levels competitive with enrollment-based approaches, without the need for pre-recordings. We further demonstrate the potential of the proposed approach for processing meeting-like recordings, where speaker activity obtained from a diarization system is used as a speaker clue for ADEnet. We show that this simple yet practical approach can successfully extract speakers after diarization, which leads to improved ASR performance when using a single microphone, especially in high overlapping conditions, with a relative word error rate reduction of up to 25 %.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech ExtractionSimilar Papers 제목 키워드 기반
Learning-based Robust Speaker Counting and Separation with the Aid of Spatial Coherence
A three-stage approach is proposed for speaker counting and speech separation in noisy and reverberant environments. In the spatial feature extraction, a spatial coherence matrix (SCM) is computed using whitened relative…
Speaker SeparationSpeech SeparationNeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection
Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recor…
EEGPlug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co…
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
Accurately predicting heart activity and other biological signals is crucial for diagnosis and monitoring. Given that speech is an outcome of multiple physiological systems, a significant body of work studied the acousti…
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a spe…
Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1