paper-with-me

홈 › Papers

A Comparative Re-Assessment of Feature Extractors for Deep Speaker Embeddings

2020-07-30 · Xuechen Liu, Md Sahidullah, Tomi Kinnunen

Modern automatic speaker verification relies largely on deep neural networks (DNNs) trained on mel-frequency cepstral coefficient (MFCC) features. While there are alternative feature extraction methods based on phase, prosody and long-term temporal operations, they have not been extensively studied with DNN-based methods. We aim to fill this gap by providing extensive re-assessment of 14 feature extractors on VoxCeleb and SITW datasets. Our findings reveal that features equipped with techniques such as spectral centroids, group delay function, and integrated noise suppression provide promising alternatives to MFCCs for deep speaker embeddings extraction. Experimental results demonstrate up to 16.3\% (VoxCeleb) and 25.1\% (SITW) relative decrease in equal error rate (EER) to the baseline.

📄 PDF Abstract BibTeX arXiv:2007.15283

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

Separating Content from Speaker Identity in Speech for the Assessment of Cognitive Impairments

2022-03-21 · Dongseok Heo, Cheul Young Park, Jaemin Cheun, Myung Jin Ko

Deep speaker embeddings have been shown effective for assessing cognitive impairments aside from their original purpose of speaker verification. However, the research found that speaker embeddings encode speaker identity…

Speaker VerificationVoice Conversion

STC speaker recognition systems for the NIST SRE 2021

2021-11-03 · Anastasia Avdeeva, Aleksei Gusev, Igor Korsunov, Alexander Kozlov 외

This paper presents a description of STC Ltd. systems submitted to the NIST 2021 Speaker Recognition Evaluation for both fixed and open training conditions. These systems consists of a number of diverse subsystems based …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker RecognitionSpeaker Verification+2

State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data

2024-10-03 · Sara Barahona, Ladislav Mošner, Themos Stafylakis, Oldřich Plchot 외

In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebr…

Speaker Verification

A Comparative Study of Pre-trained Speech and Audio Embeddings for Speech Emotion Recognition

2023-04-22 · Orchid Chetia Phukan, Arun Balaji Buduru, Rajesh Sharma

Pre-trained models (PTMs) have shown great promise in the speech and audio domain. Embeddings leveraged from these models serve as inputs for learning algorithms with applications in various downstream tasks. One such cr…

Emotion RecognitionSpeaker RecognitionSpeech Emotion Recognition

Towards Speech-only Opinion-level Sentiment Analysis

2022-06-01 · LREC 2022 6 · Annalena Aicher, Alisa Gazizullina, Aleksei Gusev, Yuri Matveev 외

The growing popularity of various forms of Spoken Dialogue Systems (SDS) raises the demand for their capability of implicitly assessing the speaker’s sentiment from speech only. Mapping the latter on user preferences ena…

Sentiment AnalysisSpeaker VerificationSpoken Dialogue Systems