Multi-scale speaker embedding-based graph attention networks for speaker diarisation
The objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker segments and the discriminative power of the embedding, according to the segment length used for embedding extraction. To this end, recent works have proposed the use of multi-scale embeddings where segments with varying lengths are used. However, the scores are combined using a weighted summation scheme where the weights are fixed after the training phase, whereas the importance of segment lengths can differ with in a single session. To address this issue, we present three key contributions in this paper: (1) we propose graph attention networks for multi-scale speaker diarisation; (2) we design scale indicators to utilise scale information of each embedding; (3) we adapt the attention-based aggregation to utilise a pre-computed affinity matrix from multi-scale embeddings. We demonstrate the effectiveness of our method in various datasets where the speaker confusion which constitutes the primary metric drops over 10% in average relative compared to the baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
Graph AttentionSimilar Papers 제목 키워드 기반
Graph Attention Networks for Speaker Verification
This work presents a novel back-end framework for speaker verification using graph attention networks. Segment-wise speaker embeddings extracted from multiple crops within an utterance are interpreted as node representat…
Graph AttentionSpeaker VerificationSpEx: Multi-Scale Time Domain Speaker Extraction Network
Speaker extraction aims to mimic humans' selective auditory attention by extracting a target speaker's voice from a multi-talker environment. It is common to perform the extraction in frequency-domain, and reconstruct th…
DecoderMulti-Task LearningSpeaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
In speaker diarization, traditional clustering-based methods remain widely used in real-world applications. However, these methods struggle with the complex distribution of speaker embeddings and overlapping speech segme…
Action DetectionActivity DetectionClusteringCommunity Detection+3Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not only for speaker recognition but also for v…
speaker-diarizationSpeaker DiarizationSpeaker RecognitionSpeaker VerificationSemi-supervised multi-channel speaker diarization with cross-channel attention
Most neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised speaker diarization system to utilize la…
speaker-diarizationSpeaker Diarization