paper-with-me

Papers

ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency

2024-06-04 · Yafeng Chen, Siqi Zheng, Hui Wang, Luyao Cheng, Qian Chen, Shiliang Zhang, Junjie Li

Speaker verification systems experience significant performance degradation when tasked with short-duration trial recordings. To address this challenge, a multi-scale feature fusion approach has been proposed to effectively capture speaker characteristics from short utterances. Constrained by the model's size, a robust backbone Enhanced Res2Net (ERes2Net) combining global and local feature fusion demonstrates sub-optimal performance in short-duration speaker verification. To further improve the short-duration feature extraction capability of ERes2Net, we expand the channel dimension within each stage. However, this modification also increases the number of model parameters and computational complexity. To alleviate this problem, we propose an improved ERes2NetV2 by pruning redundant structures, ultimately reducing both the model parameters and its computational cost. A range of experiments conducted on the VoxCeleb datasets exhibits the superiority of ERes2NetV2, which achieves EER of 0.61% for the full-duration trial, 0.98% for the 3s-duration trial, and 1.48% for the 2s-duration trial on VoxCeleb1-O, respectively.

📄 PDF Abstract BibTeX arXiv:2406.02167

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencySpeaker Verification

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Residual Connection 설명 없음
Res2Net Block A Res2Net Block is an image model block that constructs hierarchical residual-like connections within one single [residual…
Kaiming Initialization 설명 없음
Average Pooling 설명 없음
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Batch Normalization 설명 없음

Similar Papers 제목 키워드 기반

Utterance partitioning for speaker recognition: an experimental review and analysis with new findings under GMM-SVM framework

2021-05-25 · Nirmalya Sen, Md Sahidullah, Hemant Patil, Shyamal Kumar Das Mandal 외

The performance of speaker recognition system is highly dependent on the amount of speech used in enrollment and test. This work presents a detailed experimental review and analysis of the GMM-SVM based speaker recogniti…

Speaker Recognition

A Deep Neural Network for Short-Segment Speaker Recognition

2019-07-22 · Amirhossein Hajavi, Ali Etemad

Todays interactive devices such as smart-phone assistants and smart speakers often deal with short-duration speech segments. As a result, speaker recognition systems integrated into such devices will be much better suite…

Speaker Recognition

Short-duration Speaker Verification (SdSV) Challenge 2021: the Challenge Evaluation Plan

2019-12-13 · Hossein Zeinali, Kong Aik Lee, Jahangir Alam, Lukas Burget

This document describes the Short-duration Speaker Verification (SdSV) Challenge 2021. The main goal of the challenge is to evaluate new technologies for text-dependent (TD) and text-independent (TI) speaker verification…

Speaker RecognitionSpeaker VerificationText-Dependent Speaker VerificationText-Independent Speaker Verification

DAME: Duration-Aware Matryoshka Embedding for Duration-Robust Speaker Verification

2026-01-20 · Youngmoon Jung, Joon-Young Yang, Ju-ho Kim, Jaeyoung Roh 외 arxiv

Short-utterance speaker verification remains challenging due to limited speaker-discriminative cues in short speech segments. While existing methods focus on enhancing speaker encoders, the embedding learning strategy st…

Speaker Verification

Deep Speaker Embeddings for Far-Field Speaker Recognition on Short Utterances

2020-02-14 · Aleksei Gusev, Vladimir Volokhov, Tseren Andzhukaev, Sergey Novoselov 외

Speaker recognition systems based on deep speaker embeddings have achieved significant performance in controlled conditions according to the results obtained for early NIST SRE (Speaker Recognition Evaluation) datasets. …

Speaker RecognitionSpeaker Verification