From Dialect Gaps to Identity Maps: Tackling Variability in Speaker Verification
The complexity and difficulties of Kurdish speaker detection among its several dialects are investigated in this work. Because of its great phonetic and lexical differences, Kurdish with several dialects including Kurmanji, Sorani, and Hawrami offers special challenges for speaker recognition systems. The main difficulties in building a strong speaker identification system capable of precisely identifying speakers across several dialects are investigated in this work. To raise the accuracy and dependability of these systems, it also suggests solutions like sophisticated machine learning approaches, data augmentation tactics, and the building of thorough dialect-specific corpus. The results show that customized strategies for every dialect together with cross-dialect training greatly enhance recognition performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSpeaker IdentificationSpeaker RecognitionSpeaker VerificationSimilar Papers 제목 키워드 기반
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
Vietnamese exhibits substantial dialectal phonetic variation across Northern, Central, and Southern regions, where identical lexical items may be realized with markedly different pronunciations. Such variation poses chal…
Speech RecognitionDialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles
As state-of-the-art Large Language Models (LLMs) have become ubiquitous, ensuring equitable performance across diverse demographics is critical. However, it remains unclear whether these disparities arise from the explic…
Semantic SimilarityComparing Methods for Measuring Dialect Similarity in Norwegian
The present article presents four experiments with two different methods for measuring dialect similarity in Norwegian: the Levenshtein method and the neural long short term memory (LSTM) autoencoder network, a machine l…
BIG-bench Machine LearningLearning about Spanish dialects through Twitter
This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. Th…
BIG-bench Machine LearningDialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
There has been little systematic study on how dialectal differences affect toxicity detection by modern LLMs. Furthermore, although using LLMs as evaluators ("LLM-as-a-judge") is a growing research area, their sensitivit…