paper-with-me

홈 › Papers

Centroid-based deep metric learning for speaker recognition

2019-02-06 · Jixuan Wang, Kuan-Chieh Wang, Marc Law, Frank Rudzicz, Michael Brudno

Speaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant performance gap between recognizing speakers in the training set and unseen speakers. The latter case corresponds to the few-shot learning task, where a trained model is evaluated on unseen classes. Here, we optimize a speaker embedding model with prototypical network loss (PNL), a state-of-the-art approach for the few-shot image classification task. The resulting embedding model outperforms the state-of-the-art triplet loss based models in both speaker verification and identification tasks, for both seen and unseen speakers.

📄 PDF Abstract BibTeX arXiv:1902.02375

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Image ClassificationFew-Shot LearningGeneral Classificationimage-classificationImage ClassificationMetric LearningSpeaker RecognitionSpeaker VerificationTriplet

Methods 이 논문이 사용한 방법론

Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

Frequency-centroid features for word recognition of non-native English speakers

2022-06-14 · Pierre Berjon, Rajib Sharma, Avishek Nag, Soumyabrata Dev

The objective of this work is to investigate complementary features which can aid the quintessential Mel frequency cepstral coefficients (MFCCs) in the task of closed, limited set word recognition for non-native English …

Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks

2025-03-07 · Youness Atif

This paper introduces a novel audio-to-image encoding framework that integrates multiple dimensions of voice characteristics into a single RGB image for speaker recognition. In this method, the green channel encodes raw …

Speaker Recognition

Absolute decision corrupts absolutely: conservative online speaker diarisation

2022-11-09 · Youngki Kwon, Hee-Soo Heo, Bong-Jin Lee, You Jin Kim 외

Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few…

Visual Recognition with Deep Nearest Centroids

2022-09-15 · Wenguan Wang, Cheng Han, Tianfei Zhou, Dongfang Liu

We devise deep nearest centroids (DNC), a conceptually elegant yet surprisingly effective network for large-scale visual recognition, by revisiting Nearest Centroids, one of the most classic and simple classifiers. Curre…

Decision Makingimage-classificationImage ClassificationObject Recognition

Length- and Noise-aware Training Techniques for Short-utterance Speaker Recognition

2020-08-27

Speaker recognition performance has been greatly improved with the emergence of deep learning. Deep neural networks show the capacity to effectively deal with impacts of noise and reverberation, making them attractive to…

Representation LearningSpeaker Recognition