Speaker Separation Using Speaker Inventories and Estimated Speech
We propose speaker separation using speaker inventories and estimated speech (SSUSIES), a framework leveraging speaker profiles and estimated speech for speaker separation. SSUSIES contains two methods, speaker separation using speaker inventories (SSUSI) and speaker separation using estimated speech (SSUES). SSUSI performs speaker separation with the help of speaker inventory. By combining the advantages of permutation invariant training (PIT) and speech extraction, SSUSI significantly outperforms conventional approaches. SSUES is a widely applicable technique that can substantially improve speaker separation performance using the output of first-pass separation. We evaluate the models on both speaker separation and speech recognition metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker SeparationSpeech Extractionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Multi-channel Conversational Speaker Separation via Neural Diarization
When dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or mee…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognition+1Online speaker diarization of meetings guided by speech separation
Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation mode…
Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1Deep Ad-hoc Beamforming Based on Speaker Extraction for Target-Dependent Speech Separation
Recently, the research on ad-hoc microphone arrays with deep learning has drawn much attention, especially in speech enhancement and separation. Because an ad-hoc microphone array may cover such a large area that multipl…
channel selectionDeep LearningSpeech EnhancementSpeech SeparationLearning-based Robust Speaker Counting and Separation with the Aid of Spatial Coherence
A three-stage approach is proposed for speaker counting and speech separation in noisy and reverberant environments. In the spatial feature extraction, a spatial coherence matrix (SCM) is computed using whitened relative…
Speaker SeparationSpeech SeparationSimultaneous Speech Extraction for Multiple Target Speakers under the Meeting Scenarios
The common target speech separation directly estimate the target source, ignoring the interrelationship between different speakers at each frame. We propose a multiple-target speech separation model (MTSS) to simultaneou…
Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+2