paper-with-me

Papers

Study on Inter and Intra Speaker Variability in Speaker Recognition

2024-11-12 · Anton Okhotnikov, Nikita Torgashov, Ivan Yakovlev, Pavel Malov, Rostislav Makarov

Optimization of a trade-off between the number of speakers and their temporal variability (or session diversity) is crucial for the development of a speaker recognition system together with making the data collection process feasible from a time perspective. In this article, we provide the analysis of dependency between inter and intra speaker variability in training data for the modern neural network-based speaker recognition system using the VoxTube dataset for text-independent speaker recognition task. Besides, an auxiliary contribution of this work is a release of upload date metadata per utterance in a VoxTube dataset. We want this article to contribute to guidelines and best practices for collecting and filtering data from media hosting platforms to facilitate the efforts of researchers in developing speaker recognition systems.

📄 PDF Abstract BibTeX arXiv:2411.07754

Code (0)

등록된 구현이 없습니다.

Tasks

DiversitySpeaker RecognitionText-Independent Speaker Recognition

Similar Papers 제목 키워드 기반

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

2025-09-18 · Miseul Kim, Soo Jin Park, Kyungguen Byun, Hyeon-Kyeong Shin 외 arxiv

Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different indi…

Speaker Diarization

The IDLAB VoxCeleb Speaker Recognition Challenge 2021 System Description

2021-09-09 · Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck

This technical report describes the IDLab submission for track 1 and 2 of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). This speaker verification competition focuses on short duration test recordings and c…

Speaker RecognitionSpeaker Verification

Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization

2024-03-24 · Linzhi Wu, Xingyu Zhang, Yakun Zhang, Changyan Zheng 외

Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading sys…

Lip Reading

Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information

2021-10-18 · Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck

This paper contains a post-challenge performance analysis on cross-lingual speaker verification of the IDLab submission to the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). We show that current speaker embeddi…

Language IdentificationSpeaker RecognitionSpeaker Verification

Phone Duration Modeling for Speaker Age Estimation in Children

2021-09-03 · Prashanth Gurunath Shivakumar, Somer Bishop, Catherine Lord, Shrikanth Narayanan

Automatic inference of important paralinguistic information such as age from speech is an important area of research with numerous spoken language technology based applications. Speaker age estimation has applications in…

Age Estimationregression