Explainable Attribute-Based Speaker Verification
This paper proposes a fully explainable approach to speaker verification (SV), a task that fundamentally relies on individual speaker characteristics. The opaque use of speaker attributes in current SV systems raises concerns of trust. Addressing this, we propose an attribute-based explainable SV system that identifies speakers by comparing personal attributes such as gender, nationality, and age extracted automatically from voice recordings. We believe this approach better aligns with human reasoning, making it more understandable than traditional methods. Evaluated on the Voxceleb1 test set, the best performance of our system is comparable with the ground truth established when using all correct attributes, proving its efficacy. Whilst our approach sacrifices some performance compared to non-explainable methods, we believe that it moves us closer to the goal of transparent, interpretable AI and lays the groundwork for future enhancements through attribute expansion.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeSpeaker VerificationSimilar Papers 제목 키워드 기반
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
In speaker verification, we use computational method to verify if an utterance matches the identity of an enrolled speaker. This task is similar to the manual task of forensic voice comparison, where linguistic analysis …
Speaker VerificationExploring Universal Speech Attributes for Speaker Verification with an Improved Cross-stitch Network
The universal speech attributes for x-vector based speaker verification (SV) are addressed in this paper. The manner and place of articulation form the fundamental speech attribute unit (SAU), and then new speech attribu…
AttributeSpeaker VerificationFrom Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback Constraint
High-fidelity speech can be synthesized by end-to-end text-to-speech models in recent years. However, accessing and controlling speech attributes such as speaker identity, prosody, and emotion in a text-to-speech system …
Speaker VerificationSpeech Synthesistext-to-speechText to Speech+1Leveraging speaker attribute information using multi task learning for speaker verification and diarization
Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the mul…
AttributeMulti-Task LearningSpeaker RecognitionSpeaker VerificationAnalyzing speaker verification embedding extractors and back-ends under language and channel mismatch
In this paper, we analyze the behavior and performance of speaker embeddings and the back-end scoring model under domain and language mismatch. We present our findings regarding ResNet-based speaker embedding architectur…
AttributeSpeaker Verification