paper-with-me

Papers

End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

2021-05-05 · Soumi Maiti, Hakan Erdogan, Kevin Wilson, Scott Wisdom, Shinji Watanabe, John R. Hershey

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforward handling of discriminative training, unlike traditional clustering-based diarization methods. The proposed system is designed to handle meetings with unknown numbers of speakers, using variable-number permutation-invariant cross-entropy based loss functions. We introduce several components that appear to help with diarization performance, including a local convolutional network followed by a global self-attention module, multi-task transfer learning using a speaker identification component, and a sequential approach where the model is refined with a second stage. These are trained and validated on simulated meeting data based on LibriSpeech and LibriTTS datasets; final evaluations are done using LibriCSS, which consists of simulated meetings recorded using real acoustics via loudspeaker playback. The proposed model performs better than previously proposed end-to-end diarization models on these data.

📄 PDF Abstract BibTeX arXiv:2105.02096

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringSpeaker IdentificationTransfer Learning

Similar Papers 제목 키워드 기반

Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors

2022-06-06 · Shota Horiguchi, Shinji Watanabe, Paola Garcia, Yuki Takashima 외

A method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulatin…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONspeaker-diarizationSpeaker Diarization

Neural Speaker Diarization with Speaker-Wise Chain Rule

2020-06-02 · Yusuke Fujita, Shinji Watanabe, Shota Horiguchi, Yawen Xue 외

Speaker diarization is an essential step for processing multi-speaker audio. Although an end-to-end neural diarization (EEND) method achieved state-of-the-art performance, it is limited to a fixed number of speakers. In …

speaker-diarizationSpeaker Diarization

Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors

2021-07-04 · Shota Horiguchi, Shinji Watanabe, Paola Garcia, Yawen Xue 외

Attractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the cas…

Clustering

Utterance-by-utterance overlap-aware neural diarization with Graph-PIT

2022-07-28 · Keisuke Kinoshita, Thilo von Neumann, Marc Delcroix, Christoph Boeddeker 외

Recent speaker diarization studies showed that integration of end-to-end neural diarization (EEND) and clustering-based diarization is a promising approach for achieving state-of-the-art performance on various tasks. Suc…

ClusteringSegmentationspeaker-diarizationSpeaker Diarization

BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers

2020-11-05 · Eunjung Han, Chul Lee, Andreas Stolcke

We present a novel online end-to-end neural diarization system, BW-EDA-EEND, that processes data incrementally for a variable number of speakers. The system is based on the Encoder-Decoder-Attractor (EDA) architecture of…

ClusteringDecoderspeaker-diarizationSpeaker Diarization