Speaker embeddings by modeling channel-wise correlations
Speaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a linear layer, using either average or attentive pooling along the time axis. In this paper we examine an alternative pooling method, where pairwise correlations between channels for given frequencies are used as statistics. The method is inspired by style-transfer methods in computer vision, where the style of an image, modeled by the matrix of channel-wise correlations, is transferred to another image, in order to produce a new image having the style of the first and the content of the second. By drawing analogies between image style and speaker characteristics, and between image content and phonetic sequence, we explore the use of such channel-wise correlations features to train a ResNet architecture in an end-to-end fashion. Our experiments on VoxCeleb demonstrate the effectiveness of the proposed pooling method in speaker recognition.
Code (1)
Tasks
Speaker RecognitionStyle TransferMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Modeling Speaker-Listener Interaction for Backchannel Prediction
We present our latest findings on backchannel modeling novelly motivated by the canonical use of the minimal responses Yeah and Uh-huh in English and their correspondent tokens in German, and the effect of encoding the s…
PredictionSpeaker Clustering in Textual Dialogue with Utterance Correlation and Cross-corpus Dialogue Act Supervision
We propose a textual dialogue speaker clustering model, which groups the utterances of a multi-party dialogue without speaker annotations, so that the real speakers are identical inside each cluster. We find that, even w…
ClusteringCross-corpusDialogue Act ClassificationLanguage Modeling+1Improving Transformer-based Networks With Locality For Automatic Speaker Verification
Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embed…
Speaker VerificationExtracting speaker and emotion information from self-supervised speech models via channel-wise correlations
Self-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typ…
DescriptiveSelf-Supervised LearningSpatial-aware Speaker Diarization for Multi-channel Multi-party Meeting
This paper describes a spatial-aware speaker diarization system for the multi-channel multi-party meeting. The diarization system obtains direction information of speaker by microphone array. Speaker spatial embedding is…
speaker-diarizationSpeaker Diarization