paper-with-me

Papers

Deep Neural Network Embeddings with Gating Mechanisms for Text-Independent Speaker Verification

2019-03-28 · Lanhua You, Wu Guo, Li-Rong Dai, Jun Du

In this paper, gating mechanisms are applied in deep neural network (DNN) training for x-vector-based text-independent speaker verification. First, a gated convolution neural network (GCNN) is employed for modeling the frame-level embedding layers. Compared with the time-delay DNN (TDNN), the GCNN can obtain more expressive frame-level representations through carefully designed memory cell and gating mechanisms. Moreover, we propose a novel gated-attention statistics pooling strategy in which the attention scores are shared with the output gate. The gated-attention statistics pooling combines both gating and attention mechanisms into one framework; therefore, we can capture more useful information in the temporal pooling layer. Experiments are carried out using the NIST SRE16 and SRE18 evaluation datasets. The results demonstrate the effectiveness of the GCNN and show that the proposed gated-attention statistics pooling can further improve the performance.

📄 PDF Abstract BibTeX arXiv:1903.12092

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationText-Independent Speaker Verification

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model

2018-09-12 · Suwon Shon, Hao Tang, James Glass

In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding …

Speaker RecognitionText-Independent Speaker Recognition

Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification

2019-06-19 · Youngmoon Jung, Younggwan Kim, Hyungjun Lim, Yeunju Choi 외

In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification. We first partition the output feature maps from a deep residu…

Speaker VerificationText-Independent Speaker Verification

Attention-based conditioning methods using variable frame rate for style-robust speaker verification

2022-06-28 · Amber Afshan, Abeer Alwan

We propose an approach to extract speaker embeddings that are robust to speaking style variations in text-independent speaker verification. Typically, speaker embedding extraction includes training a DNN for speaker clas…

Speaker VerificationText-Independent Speaker Verification

Auditory Separation of a Conversation from Background via Attentional Gating

2019-05-26 · Shariq Mobin, Bruno Olshausen

We present a model for separating a set of voices out of a sound mixture containing an unknown number of sources. Our Attentional Gating Network (AGN) uses a variable attentional context to specify which speakers in the …

Speaker Separation

Triplet Based Embedding Distance and Similarity Learning for Text-independent Speaker Verification

2019-08-06 · Zongze Ren, Zhiyong Chen, Shugong Xu

Speaker embeddings become growing popular in the text-independent speaker verification task. In this paper, we propose two improvements during the training stage. The improvements are both based on triplet cause the trai…

Speaker RecognitionSpeaker VerificationText-Independent Speaker VerificationTriplet