ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
Speaker verification systems experience significant performance degradation when tasked with short-duration trial recordings. To address this challenge, a multi-scale feature fusion approach has been proposed to effectively capture speaker characteristics from short utterances. Constrained by the model's size, a robust backbone Enhanced Res2Net (ERes2Net) combining global and local feature fusion demonstrates sub-optimal performance in short-duration speaker verification. To further improve the short-duration feature extraction capability of ERes2Net, we expand the channel dimension within each stage. However, this modification also increases the number of model parameters and computational complexity. To alleviate this problem, we propose an improved ERes2NetV2 by pruning redundant structures, ultimately reducing both the model parameters and its computational cost. A range of experiments conducted on the VoxCeleb datasets exhibits the superiority of ERes2NetV2, which achieves EER of 0.61% for the full-duration trial, 0.98% for the 3s-duration trial, and 1.48% for the 2s-duration trial on VoxCeleb1-O, respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
Computational EfficiencySpeaker VerificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Utterance partitioning for speaker recognition: an experimental review and analysis with new findings under GMM-SVM framework
The performance of speaker recognition system is highly dependent on the amount of speech used in enrollment and test. This work presents a detailed experimental review and analysis of the GMM-SVM based speaker recogniti…
Speaker RecognitionA Deep Neural Network for Short-Segment Speaker Recognition
Todays interactive devices such as smart-phone assistants and smart speakers often deal with short-duration speech segments. As a result, speaker recognition systems integrated into such devices will be much better suite…
Speaker RecognitionShort-duration Speaker Verification (SdSV) Challenge 2021: the Challenge Evaluation Plan
This document describes the Short-duration Speaker Verification (SdSV) Challenge 2021. The main goal of the challenge is to evaluate new technologies for text-dependent (TD) and text-independent (TI) speaker verification…
Speaker RecognitionSpeaker VerificationText-Dependent Speaker VerificationText-Independent Speaker VerificationDAME: Duration-Aware Matryoshka Embedding for Duration-Robust Speaker Verification
Short-utterance speaker verification remains challenging due to limited speaker-discriminative cues in short speech segments. While existing methods focus on enhancing speaker encoders, the embedding learning strategy st…
Speaker VerificationDeep Speaker Embeddings for Far-Field Speaker Recognition on Short Utterances
Speaker recognition systems based on deep speaker embeddings have achieved significant performance in controlled conditions according to the results obtained for early NIST SRE (Speaker Recognition Evaluation) datasets. …
Speaker RecognitionSpeaker Verification