paper-with-me

Papers

Speaker embeddings by modeling channel-wise correlations

2021-04-06 · Themos Stafylakis, Johan Rohdin, Lukas Burget

Speaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a linear layer, using either average or attentive pooling along the time axis. In this paper we examine an alternative pooling method, where pairwise correlations between channels for given frequencies are used as statistics. The method is inspired by style-transfer methods in computer vision, where the style of an image, modeled by the matrix of channel-wise correlations, is transferred to another image, in order to produce a new image having the style of the first and the content of the second. By drawing analogies between image style and speaker characteristics, and between image content and phonetic sequence, we explore the use of such channel-wise correlations features to train a ResNet architecture in an end-to-end fashion. Our experiments on VoxCeleb demonstrate the effectiveness of the proposed pooling method in speaker recognition.

📄 PDF Abstract BibTeX arXiv:2104.02571

Code (1)

tstafylakis/Speaker-Embeddings-Correlation-Pooling 공식 구현 tf

Tasks

Speaker RecognitionStyle Transfer

Methods 이 논문이 사용한 방법론

Residual Connection 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Average Pooling 설명 없음
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…
Batch Normalization 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Bottleneck Residual Block A Bottleneck Residual Block is a variant of the residual block that utilises 1x1 convolutions to create a bottleneck. The…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…

Similar Papers 제목 키워드 기반

Modeling Speaker-Listener Interaction for Backchannel Prediction

2023-04-10 · Daniel Ortega, Sarina Meyer, Antje Schweitzer, Ngoc Thang Vu

We present our latest findings on backchannel modeling novelly motivated by the canonical use of the minimal responses Yeah and Uh-huh in English and their correspondent tokens in German, and the effect of encoding the s…

Prediction

Speaker Clustering in Textual Dialogue with Utterance Correlation and Cross-corpus Dialogue Act Supervision

2022-01-16 · ACL ARR January 2022 1 · Anonymous

We propose a textual dialogue speaker clustering model, which groups the utterances of a multi-party dialogue without speaker annotations, so that the real speakers are identical inside each cluster. We find that, even w…

ClusteringCross-corpusDialogue Act ClassificationLanguage Modeling+1

Improving Transformer-based Networks With Locality For Automatic Speaker Verification

2023-02-17 · Mufan Sang, Yong Zhao, Gang Liu, John H. L. Hansen 외

Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embed…

Speaker Verification

Extracting speaker and emotion information from self-supervised speech models via channel-wise correlations

2022-10-15 · Themos Stafylakis, Ladislav Mosner, Sofoklis Kakouros, Oldrich Plchot 외

Self-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typ…

DescriptiveSelf-Supervised Learning

Spatial-aware Speaker Diarization for Multi-channel Multi-party Meeting

2022-09-24 · Jie Wang, Yuji Liu, Binling Wang, Yiming Zhi 외

This paper describes a spatial-aware speaker diarization system for the multi-channel multi-party meeting. The diarization system obtains direction information of speaker by microphone array. Speaker spatial embedding is…

speaker-diarizationSpeaker Diarization