paper-with-me

Papers

RawNeXt: Speaker verification system for variable-duration utterances with deep layer aggregation and extended dynamic scaling policies

2021-12-15 · Ju-ho Kim, Hye-jin Shim, Jungwoo Heo, Ha-Jin Yu

Despite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker verification system called RawNeXt that can handle input raw waveforms of arbitrary length by employing the following two components: (1) A deep layer aggregation strategy enhances speaker information by iteratively and hierarchically aggregating features of various time scales and spectral channels output from blocks. (2) An extended dynamic scaling policy flexibly processes features according to the length of the utterance by selectively merging the activations of different resolution branches in each block. Owing to these two components, our proposed model can extract speaker embeddings rich in time-spectral information and operate dynamically on length variations. Experimental results on the VoxCeleb1 test set consisting of various duration utterances demonstrate that RawNeXt achieves state-of-the-art performance compared to the recently proposed systems. Our code and trained model weights are available at https://github.com/wngh1187/RawNeXt.

📄 PDF Abstract BibTeX arXiv:2112.07935

Code (1)

wngh1187/rawnext 공식 구현 pytorch

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

MR-RawNet: Speaker verification system with multiple temporal resolutions for variable duration utterances using raw waveforms

2024-06-11 · Seung-bin Kim, Chan-yeong Lim, Jungwoo Heo, Ju-ho Kim 외

In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To…

Speaker Verification

Improving Multi-Scale Aggregation Using Feature Pyramid Module for Robust Speaker Verification of Variable-Duration Utterances

2020-04-07 · Youngmoon Jung, Seong Min Kye, Yeunju Choi, Myunghun Jung 외

Currently, the most widely used approach for speaker verification is the deep speaker embedding learning. In this approach, we obtain a speaker embedding vector by pooling single-scale features that are extracted from th…

Speaker VerificationText-Independent Speaker Verification

Short-duration Speaker Verification (SdSV) Challenge 2021: the Challenge Evaluation Plan

2019-12-13 · Hossein Zeinali, Kong Aik Lee, Jahangir Alam, Lukas Burget

This document describes the Short-duration Speaker Verification (SdSV) Challenge 2021. The main goal of the challenge is to evaluate new technologies for text-dependent (TD) and text-independent (TI) speaker verification…

Speaker RecognitionSpeaker VerificationText-Dependent Speaker VerificationText-Independent Speaker Verification

Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization

2024-12-22 · Natalia Tomashenko, Emmanuel Vincent, Marc Tommasi

In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verific…

Speaker Verification

Universal speaker recognition encoders for different speech segments duration

2022-10-28 · Sergey Novoselov, Vladimir Volokhov, Galina Lavrentyeva

Creating universal speaker encoders which are robust for different acoustic and speech duration conditions is a big challenge today. According to our observations systems trained on short speech segments are optimal for …

Speaker RecognitionSpeaker Verification