paper-with-me

홈 › Papers

Magnitude-aware Probabilistic Speaker Embeddings

2022-02-28 · Nikita Kuzmin, Igor Fedorov, Alexey Sholokhov

Recently, hyperspherical embeddings have established themselves as a dominant technique for face and voice recognition. Specifically, Euclidean space vector embeddings are learned to encode person-specific information in their direction while ignoring the magnitude. However, recent studies have shown that the magnitudes of the embeddings extracted by deep neural networks may indicate the quality of the corresponding inputs. This paper explores the properties of the magnitudes of the embeddings related to quality assessment and out-of-distribution detection. We propose a new probabilistic speaker embedding extractor using the information encoded in the embedding magnitude and leverage it in the speaker verification pipeline. We also propose several quality-aware diarization methods and incorporate the magnitudes in those. Our results indicate significant improvements over magnitude-agnostic baselines both in speaker verification and diarization tasks.

📄 PDF Abstract BibTeX arXiv:2202.13826

Code (1)

clovaai/voxceleb_trainer 공식 구현 pytorch

Tasks

Out-of-Distribution DetectionSpeaker Verification

Similar Papers 제목 키워드 기반

Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification

2021-12-16 · Sung Hwan Mun, Min Hyun Han, Dongjune Lee, JiHwan Kim 외

In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and an uncertainty-aware probabilistic spea…

Contrastive LearningRepresentation LearningSpeaker Verification

Probabilistic embeddings for speaker diarization

2020-04-06 · Anna Silnova, Niko Brümmer, Johan Rohdin, Themos Stafylakis 외

Speaker embeddings (x-vectors) extracted from very short segments of speech have recently been shown to give competitive performance in speaker diarization. We generalize this recipe by extracting from each speech segmen…

Clusteringspeaker-diarizationSpeaker Diarization

Speaker Diarization using Deep Recurrent Convolutional Neural Networks for Speaker Embeddings

2017-08-09 · Pawel Cyrta, Tomasz Trzciński, Wojciech Stokowiec

In this paper we propose a new method of speaker diarization that employs a deep learning architecture to learn speaker embeddings. In contrast to the traditional approaches that build their speaker embeddings using manu…

speaker-diarizationSpeaker Diarization

Content-Aware Speaker Embeddings for Speaker Diarisation

2021-02-12 · G. Sun, D. Liu, C. Zhang, P. C. Woodland

Recent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddings. In this paper, the content-aware sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringSpeaker Recognition+3

Multi-Level Speaker Representation for Target Speaker Extraction

2024-10-21 · Ke Zhang, Junjie Li, Shuai Wang, Yangjie Wei 외

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with…

Target Speaker Extraction