paper-with-me

홈 › Papers

DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding

2021-05-28 · Neil Zeghidour, Olivier Teboul, David Grangier

We introduce DIVE, an end-to-end speaker diarization algorithm. Our neural algorithm presents the diarization task as an iterative process: it repeatedly builds a representation for each speaker before predicting the voice activity of each speaker conditioned on the extracted representations. This strategy intrinsically resolves the speaker ordering ambiguity without requiring the classical permutation invariant training loss. In contrast with prior work, our model does not rely on pretrained speaker representations and optimizes all parameters of the system with a multi-speaker voice activity loss. Importantly, our loss explicitly excludes unreliable speaker turn boundaries from training, which is adapted to the standard collar-based Diarization Error Rate (DER) evaluation. Overall, these contributions yield a system redefining the state-of-the-art on the standard CALLHOME benchmark, with 6.7% DER compared to 7.8% for the best alternative.

📄 PDF Abstract BibTeX arXiv:2105.13802

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

Simultaneous Speech Recognition and Speaker Diarization for Monaural Dialogue Recordings with Target-Speaker Acoustic Models

2019-09-17 · Naoyuki Kanda, Shota Horiguchi, Yusuke Fujita, Yawen Xue 외

This paper investigates the use of target-speaker automatic speech recognition (TS-ASR) for simultaneous speech recognition and speaker diarization of single-channel dialogue recordings. TS-ASR is a technique to automati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeaker-diarization+3

End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors

2020-05-20 · Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue 외

End-to-end speaker diarization for an unknown number of speakers is addressed in this paper. Recently proposed end-to-end speaker diarization outperformed conventional clustering-based speaker diarization, but it has one…

ClusteringDecoderspeaker-diarizationSpeaker Diarization

Compositional embedding models for speaker identification and diarization with simultaneous speech from 2+ speakers

2020-10-22 · Zeqian Li, Jacob Whitehill

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compos…

speaker-diarizationSpeaker DiarizationSpeaker Identification

Speaker Diarization with Lexical Information

2020-04-13 · Tae Jin Park, Kyu J. Han, Jing Huang, Xiaodong He 외

This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeaker-diarization+3

Speaker Diarization with Region Proposal Network

2020-02-14 · Zili Huang, Shinji Watanabe, Yusuke Fujita, Paola Garcia 외

Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in vario…

Region Proposalspeaker-diarizationSpeaker Diarization