Bilingual Text-dependent Speaker Verification with Pre-trained Models for TdSV Challenge 2024
This paper presents our submissions to the Iranian division of the Text-dependent Speaker Verification Challenge (TdSV) 2024. TdSV aims to determine if a specific phrase was spoken by a target speaker. We developed two independent subsystems based on pre-trained models: For phrase verification, a phrase classifier rejected incorrect phrases, while for speaker verification, a pre-trained ResNet293 with domain adaptation extracted speaker embeddings for computing cosine similarity scores. In addition, we evaluated Whisper-PMFA, a pre-trained ASR model adapted for speaker verification, and found that, although it outperforms randomly initialized ResNets, it falls short of the performance of pre-trained ResNets, highlighting the importance of large-scale pre-training. The results also demonstrate that achieving competitive performance on TdSV without joint modeling of speaker and text is possible. Our best system achieved a MinDCF of 0.0358 on the evaluation subset and won the challenge.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationSpeaker VerificationText-Dependent Speaker VerificationSimilar Papers 제목 키워드 기반
Cross-lingual Multispeaker Text-to-Speech under Limited-Data Scenario
Modeling voices for multiple speakers and multiple languages in one text-to-speech system has been a challenge for a long time. This paper presents an extension on Tacotron2 to achieve bilingual multispeaker speech synth…
AttributeSpeech Synthesistext-to-speechText to SpeechText-Independent Speaker Verification Using Long Short-Term Memory Networks
In this paper, an architecture based on Long Short-Term Memory Networks has been proposed for the text-independent scenario which is aimed to capture the temporal speaker-related information by operating over traditional…
Speaker VerificationText-Independent Speaker VerificationText-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report
This paper presents a system for the 2024 Text-Dependent Speaker Verification (TdSV) Challenge. The system achieved a Minimum Detection Cost Function (MinDCF) of 0.0461 and an Equal Error Rate (EER) of 1.3\%. Our approac…
Speaker VerificationEnsemble LearningData AugmentationText-dependent Speaker Verification (TdSV) Challenge 2024: Challenge Evaluation Plan
This document outlines the Text-dependent Speaker Verification (TdSV) Challenge 2024, which centers on analyzing and exploring novel approaches for text-dependent speaker verification. The primary goal of this challenge …
Few-Shot LearningMulti-Task LearningSelf-Supervised LearningSpeaker Verification+1Deep Speaker Vectors for Semi Text-independent Speaker Verification
Recent research shows that deep neural networks (DNNs) can be used to extract deep speaker vectors (d-vectors) that preserve speaker characteristics and can be used in speaker verification. This new method has been teste…
Speaker RecognitionSpeaker VerificationText-Dependent Speaker VerificationText-Independent Speaker Recognition+1