paper-with-me

홈 › Papers

oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models

2024-09-16 · Muhammad Sudipto Siam Dip, Md Anik Hasan, Sapnil Sarker Bipro, Md Abdur Raiyan, Mohammod Abdul Motin

In this study, we address the challenge of speaker recognition using a novel data augmentation technique of adding noise to enrollment files. This technique efficiently aligns the sources of test and enrollment files, enhancing comparability. Various pre-trained models were employed, with the resnet model achieving the highest DCF of 0.84 and an EER of 13.44. The augmentation technique notably improved these results to 0.75 DCF and 12.79 EER for the resnet model. Comparative analysis revealed the superiority of resnet over models such as ECPA, Mel-spectrogram, Payonnet, and Titanet large. Results, along with different augmentation schemes, contribute to the success of RoboVox far-field speaker recognition in this paper

📄 PDF Abstract BibTeX arXiv:2409.10240

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationSpeaker Recognition

Methods 이 논문이 사용한 방법론

Kaiming Initialization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Average Pooling 설명 없음
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…

Similar Papers 제목 키워드 기반

Team HYU ASML ROBOVOX SP Cup 2024 System Description

2024-07-16 · Jeong-Hwan Choi, Gaeun Kim, Hee-Jae Lee, Seyun Ahn 외

This report describes the submission of HYU ASML team to the IEEE Signal Processing Cup 2024 (SP Cup 2024). This challenge, titled "ROBOVOX: Far-Field Speaker Recognition by a Mobile Robot," focuses on speaker recognitio…

Data AugmentationSpeaker Recognition

STC Speaker Recognition Systems for the VOiCES From a Distance Challenge

2019-04-12 · Sergey Novoselov, Aleksei Gusev, Artem Ivanov, Timur Pekhovsky 외

This paper presents the Speech Technology Center (STC) speaker recognition (SR) systems submitted to the VOiCES From a Distance challenge 2019. The challenge's SR task is focused on the problem of speaker recognition in …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMetric Learning+4

Augmentation adversarial training for self-supervised speaker recognition

2020-07-23 · Jaesung Huh, Hee Soo Heo, Jingu Kang, Shinji Watanabe 외

The goal of this work is to train robust speaker recognition models without speaker labels. Recent works on unsupervised speaker representations are based on contrastive learning in which they encourage within-utterance …

Contrastive LearningSpeaker Recognition

Investigation of Data Augmentation Techniques for Disordered Speech Recognition

2022-01-14 · Mengzhe Geng, Xurong Xie, Shansong Liu, Jianwei Yu 외

Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disabilities, lead to the difficulty in colle…

Data Augmentationspeech-recognitionSpeech Recognition

Non-Parallel Voice Conversion for ASR Augmentation

2022-09-15 · Gary Wang, Andrew Rosenberg, Bhuvana Ramabhadran, Fadi Biadsy 외

Automatic speech recognition (ASR) needs to be robust to speaker differences. Voice Conversion (VC) modifies speaker characteristics of input speech. This is an attractive feature for ASR data augmentation. In this paper…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDiversity+3