oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
In this study, we address the challenge of speaker recognition using a novel data augmentation technique of adding noise to enrollment files. This technique efficiently aligns the sources of test and enrollment files, enhancing comparability. Various pre-trained models were employed, with the resnet model achieving the highest DCF of 0.84 and an EER of 13.44. The augmentation technique notably improved these results to 0.75 DCF and 12.79 EER for the resnet model. Comparative analysis revealed the superiority of resnet over models such as ECPA, Mel-spectrogram, Payonnet, and Titanet large. Results, along with different augmentation schemes, contribute to the success of RoboVox far-field speaker recognition in this paper
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSpeaker RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Team HYU ASML ROBOVOX SP Cup 2024 System Description
This report describes the submission of HYU ASML team to the IEEE Signal Processing Cup 2024 (SP Cup 2024). This challenge, titled "ROBOVOX: Far-Field Speaker Recognition by a Mobile Robot," focuses on speaker recognitio…
Data AugmentationSpeaker RecognitionSTC Speaker Recognition Systems for the VOiCES From a Distance Challenge
This paper presents the Speech Technology Center (STC) speaker recognition (SR) systems submitted to the VOiCES From a Distance challenge 2019. The challenge's SR task is focused on the problem of speaker recognition in …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMetric Learning+4Augmentation adversarial training for self-supervised speaker recognition
The goal of this work is to train robust speaker recognition models without speaker labels. Recent works on unsupervised speaker representations are based on contrastive learning in which they encourage within-utterance …
Contrastive LearningSpeaker RecognitionInvestigation of Data Augmentation Techniques for Disordered Speech Recognition
Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disabilities, lead to the difficulty in colle…
Data Augmentationspeech-recognitionSpeech RecognitionNon-Parallel Voice Conversion for ASR Augmentation
Automatic speech recognition (ASR) needs to be robust to speaker differences. Voice Conversion (VC) modifies speaker characteristics of input speech. This is an attractive feature for ASR data augmentation. In this paper…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDiversity+3