ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention
In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetric Cross Attention (ACA) to replace temporal pooling. ACA is able to distill large, variable-length sequences into small, fixed-sized latents by attending a small query to large key and value matrices. In ACA-Net, we build a Multi-Layer Aggregation (MLA) block using ACA to generate fixed-sized identity vectors from variable-length inputs. Through global attention, ACA-Net acts as an efficient global feature extractor that adapts to temporal variability unlike existing SV models that apply a fixed function for pooling over the temporal dimension which may obscure information about the signal's non-stationary temporal variability. Our experiments on the WSJ0-1talker show ACA-Net outperforms a strong baseline by 5\% relative improvement in EER using only 1/5 of the parameters.
Code (1)
Tasks
Speaker VerificationSimilar Papers 제목 키워드 기반
Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
Automated speaker identification (SID) is a crucial step for the personalization of a wide range of speech-enabled services. Typical SID systems use a symmetric enrollment-verification framework with a single model to de…
Speaker IdentificationSpeaker RecognitionCross-lingual Text-independent Speaker Verification using Unsupervised Adversarial Discriminative Domain Adaptation
Speaker verification systems often degrade significantly when there is a language mismatch between training and testing data. Being able to improve cross-lingual speaker verification system using unlabeled data can great…
Domain AdaptationSpeaker VerificationText-Independent Speaker VerificationConvolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification
Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving …
Speaker VerificationText-Independent Speaker VerificationSpeaker-Utterance Dual Attention for Speaker and Utterance Verification
In this paper, we study a novel technique that exploits the interaction between speaker traits and linguistic content to improve both speaker verification and utterance verification performance. We implement an idea of s…
Speaker VerificationAudio-Visual Speaker Verification via Joint Cross-Attention
Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complem…
Speaker Verification