Study on Inter and Intra Speaker Variability in Speaker Recognition
Optimization of a trade-off between the number of speakers and their temporal variability (or session diversity) is crucial for the development of a speaker recognition system together with making the data collection process feasible from a time perspective. In this article, we provide the analysis of dependency between inter and intra speaker variability in training data for the modern neural network-based speaker recognition system using the VoxTube dataset for text-independent speaker recognition task. Besides, an auxiliary contribution of this work is a release of upload date metadata per utterance in a VoxTube dataset. We want this article to contribute to guidelines and best practices for collecting and filtering data from media hosting platforms to facilitate the efforts of researchers in developing speaker recognition systems.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversitySpeaker RecognitionText-Independent Speaker RecognitionSimilar Papers 제목 키워드 기반
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different indi…
Speaker DiarizationThe IDLAB VoxCeleb Speaker Recognition Challenge 2021 System Description
This technical report describes the IDLab submission for track 1 and 2 of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). This speaker verification competition focuses on short duration test recordings and c…
Speaker RecognitionSpeaker VerificationLandmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading sys…
Lip ReadingTackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information
This paper contains a post-challenge performance analysis on cross-lingual speaker verification of the IDLab submission to the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). We show that current speaker embeddi…
Language IdentificationSpeaker RecognitionSpeaker VerificationPhone Duration Modeling for Speaker Age Estimation in Children
Automatic inference of important paralinguistic information such as age from speech is an important area of research with numerous spoken language technology based applications. Speaker age estimation has applications in…
Age Estimationregression