End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings
We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforward handling of discriminative training, unlike traditional clustering-based diarization methods. The proposed system is designed to handle meetings with unknown numbers of speakers, using variable-number permutation-invariant cross-entropy based loss functions. We introduce several components that appear to help with diarization performance, including a local convolutional network followed by a global self-attention module, multi-task transfer learning using a speaker identification component, and a sequential approach where the model is refined with a second stage. These are trained and validated on simulated meeting data based on LibriSpeech and LibriTTS datasets; final evaluations are done using LibriCSS, which consists of simulated meetings recorded using real acoustics via loudspeaker playback. The proposed model performs better than previously proposed end-to-end diarization models on these data.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringSpeaker IdentificationTransfer LearningSimilar Papers 제목 키워드 기반
Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors
A method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulatin…
Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONspeaker-diarizationSpeaker DiarizationNeural Speaker Diarization with Speaker-Wise Chain Rule
Speaker diarization is an essential step for processing multi-speaker audio. Although an end-to-end neural diarization (EEND) method achieved state-of-the-art performance, it is limited to a fixed number of speakers. In …
speaker-diarizationSpeaker DiarizationTowards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors
Attractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the cas…
ClusteringUtterance-by-utterance overlap-aware neural diarization with Graph-PIT
Recent speaker diarization studies showed that integration of end-to-end neural diarization (EEND) and clustering-based diarization is a promising approach for achieving state-of-the-art performance on various tasks. Suc…
ClusteringSegmentationspeaker-diarizationSpeaker DiarizationBW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers
We present a novel online end-to-end neural diarization system, BW-EDA-EEND, that processes data incrementally for a variable number of speakers. The system is based on the Encoder-Decoder-Attractor (EDA) architecture of…
ClusteringDecoderspeaker-diarizationSpeaker Diarization