Guided Training: A Simple Method for Single-channel Speaker Separation
Deep learning has shown a great potential for speech separation, especially for speech and non-speech separation. However, it encounters permutation problem for multi-speaker separation where both target and interference are speech. Permutation Invariant training (PIT) was proposed to solve this problem by permuting the order of the multiple speakers. Another way is to use an anchor speech, a short speech of the target speaker, to model the speaker identity. In this paper, we propose a simple strategy to train a long short-term memory (LSTM) model to solve the permutation problem in speaker separation. Specifically, we insert a short speech of target speaker at the beginning of a mixture as guide information. So, the first appearing speaker is defined as the target. Due to the powerful capability on sequence modeling, LSTM can use its memory cells to track and separate target speech from interfering speech. Experimental results show that the proposed training strategy is effective for speaker separation.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker SeparationSpeech SeparationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Beamformer-Guided Target Speaker Extraction
We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs …
Target Speaker ExtractionEnd-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis
We present an end-to-end multichannel speaker-attributed automatic speech recognition (MC-SA-ASR) system that combines a Conformer-based encoder with multi-frame crosschannel attention and a speaker-attributed Transforme…
Automatic Speech RecognitionDecoderSpeaker Identificationspeech-recognition+1End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper…
Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech-to-Text+2Mutual Learning of Single- and Multi-Channel End-to-End Neural Diarization
Due to the high performance of multi-channel speech processing, we can use the outputs from a multi-channel model as teacher labels when training a single-channel model with knowledge distillation. To the contrary, it is…
Knowledge Distillationspeaker-diarizationSpeaker DiarizationTransfer LearningMulti-Channel Speaker Verification for Single and Multi-talker Speech
To improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features. Specifically, we combine spectral, …
Action DetectionActivity DetectionSpeaker VerificationSpeech Enhancement+1