Exploring Universal Speech Attributes for Speaker Verification with an Improved Cross-stitch Network
The universal speech attributes for x-vector based speaker verification (SV) are addressed in this paper. The manner and place of articulation form the fundamental speech attribute unit (SAU), and then new speech attribute (NSA) units for acoustic modeling are generated by tied tri-SAU states. An improved cross-stitch network is adopted as a multitask learning (MTL) framework for integrating these universal speech attributes into the x-vector network training process. Experiments are conducted on common condition 5 (CC5) of the core-core and the 10 s-10 s tests of the NIST SRE10 evaluation set, and the proposed algorithm can achieve consistent improvements over the baseline x-vector on both these tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeSpeaker VerificationSimilar Papers 제목 키워드 기반
Universal speaker recognition encoders for different speech segments duration
Creating universal speaker encoders which are robust for different acoustic and speech duration conditions is a big challenge today. According to our observations systems trained on short speech segments are optimal for …
Speaker RecognitionSpeaker VerificationFrom Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback Constraint
High-fidelity speech can be synthesized by end-to-end text-to-speech models in recent years. However, accessing and controlling speech attributes such as speaker identity, prosody, and emotion in a text-to-speech system …
Speaker VerificationSpeech Synthesistext-to-speechText to Speech+1Incorporation of Speech Duration Information in Score Fusion of Speaker Recognition Systems
In recent years identity-vector (i-vector) based speaker verification (SV) systems have become very successful. Nevertheless, environmental noise and speech duration variability still have a significant effect on degradi…
Speaker RecognitionSpeaker VerificationDELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
Self-supervised speech models have achieved remarkable success on content-driven tasks, yet they remain limited in capturing speaker-discriminative features critical for verification, diarization, and profiling applicati…
Representation LearningSpeaker VerificationExploring wav2vec 2.0 on speaker verification and language identification
Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning. It follows a two-stage training process of pre-training and fine-tuning, and performs well in speech recognition tasks espe…
Language IdentificationMulti-Task LearningRepresentation LearningSpeaker Verification+3