Compositional embedding models for speaker identification and diarization with simultaneous speech from 2+ speakers
We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding models contain a function f that separates speech from different speakers. In addition, they include a composition function g to compute set-union operations in the embedding space so as to infer the set of speakers within the input audio. In an experiment on multi-person speaker identification using synthesized LibriSpeech data, the proposed method outperforms traditional embedding methods that are only trained to separate single speakers (not speaker sets). In a speaker diarization experiment on the AMI Headset Mix corpus, we achieve state-of-the-art accuracy (DER=22.93%), slightly higher than the previous best result (23.82% from [3]).
Code (1)
Tasks
speaker-diarizationSpeaker DiarizationSpeaker IdentificationSimilar Papers 제목 키워드 기반
Compositional Clustering: Applications to Multi-Label Object Recognition and Speaker Identification
We consider a novel clustering task in which clusters can have compositional relationships, e.g., one cluster contains images of rectangles, one contains images of circles, and a third (compositional) cluster contains im…
ClusteringFew-Shot LearningObject Recognitionspeaker-diarization+2Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling
Traditional speech separation and speaker diarization approaches rely on prior knowledge of target speakers or a predetermined number of participants in audio signals. To address these limitations, recent advances focus …
Speaker DiarizationSpeech SeparationSimultaneous Speech Recognition and Speaker Diarization for Monaural Dialogue Recordings with Target-Speaker Acoustic Models
This paper investigates the use of target-speaker automatic speech recognition (TS-ASR) for simultaneous speech recognition and speaker diarization of single-channel dialogue recordings. TS-ASR is a technique to automati…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeaker-diarization+3Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a spe…
Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation
This paper proposes a novel Sequence-to-Sequence Neural Diarization (S2SND) framework to perform online and offline speaker diarization. It is developed from the sequence-to-sequence architecture of our previous target-s…
Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization