``Oh, I've Heard That Before'': Modelling Own-Dialect Bias After Perceptual Learning by Weighting Training Data
Human listeners are able to quickly and robustly adapt to new accents and do so by using information about speaker{'}s identities. This paper will present experimental evidence that, even considering information about speaker{'}s identities, listeners retain a strong bias towards the acoustics of their own dialect after dialect learning. Participants{'} behaviour was accurately mimicked by a classifier which was trained on more cases from the base dialect and fewer from the target dialect. This suggests that imbalanced training data may result in automatic speech recognition errors consistent with those of speakers from populations over-represented in the training data.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
ArbDialectID at MADAR Shared Task 1: Language Modelling and Ensemble Learning for Fine Grained Arabic Dialect Identification
In this paper, we present a Dialect Identification system (ArbDialectID) that competed at Task 1 of the MADAR shared task, MADARTravel Domain Dialect Identification. We build a course and a fine-grained identification mo…
Dialect IdentificationEnsemble LearningFeature EngineeringLanguage Modelling+1Side-by-side Comparison Amplifies Dialect Bias in Language Models
Language models (LMs) can exhibit biases based on variations in their dialects, even in the absence of a dialect label, a behavior known as covert dialect bias. In this work, we quantify covert dialect bias in online dis…
Decision MakingVoices Unheard: NLP Resources and Models for Yorùbá Regional Dialects
Yor\`ub\'a an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, res…
Automatic Speech RecognitionMachine Translationspeech-recognitionSpeech Recognition+3A Spelling Correction Corpus for Multiple Arabic Dialects
Arabic dialects are the non-standard varieties of Arabic commonly spoken {--} and increasingly written on social media {--} across the Arab world. Arabic dialects do not have standard orthographies, a challenge for natur…
Spelling CorrectionText NormalizationTwitter Data Analysis: Izmir Earthquake Case
T\"urkiye is located on a fault line; earthquakes often occur on a large and small scale. There is a need for effective solutions for gathering current information during disasters. We can use social media to get insight…
ManagementPublic RelationsSentiment Analysis