paper-with-me

홈 › Papers

Corpus Development of Kiswahili Speech Recognition Test and Evaluation sets, Preemptively Mitigating Demographic Bias Through Collaboration with Linguists

2022-05-01 · ComputEL (ACL) 2022 5 · Kathleen Siminyu, Kibibi Mohamed Amran, Abdulrahman Ndegwa Karatu, Mnata Resani, Mwimbi Makobo Junior, Rebecca Ryakitimbo, Britone Mwasaru

Language technologies, particularly speech technologies, are becoming more pervasive for access to digital platforms and resources. This brings to the forefront concerns of their inclusivity, first in terms of language diversity. Additionally, research shows speech recognition to be more accurate for men than for women and more accurate for individuals younger than 30 years of age than those older. In the Global South where languages are low resource, these same issues should be taken into consideration in data collection efforts to not replicate these mistakes. It is also important to note that in varying contexts within the Global South, this work presents additional nuance and potential for bias based on accents, related dialects and variants of a language. This paper documents i) the designing and execution of a Linguists Engagement for purposes of building an inclusive Kiswahili Speech Recognition dataset, representative of the diversity among speakers of the language ii) the unexpected yet key learning in terms of socio-linguistcs which demonstrate the importance of multi-disciplinarity in teams developing datasets and NLP technologies iii) the creation of a test dataset intended to be used for evaluating the performance of Speech Recognition models on demographic groups that are likely to be underrepresented.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Diversityspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Phonemic Representation and Transcription for Speech to Text Applications for Under-resourced Indigenous African Languages: The Case of Kiswahili

2022-10-29 · Ebbie Awino, Lilian Wanzare, Lawrence Muchemi, Barack Wanjawa 외

Building automatic speech recognition (ASR) systems is a challenging task, especially for under-resourced languages that need to construct corpora nearly from scratch and lack sufficient training data. It has emerged tha…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Algorithm for Semantic Network Generation from Texts of Low Resource Languages Such as Kiswahili

2025-01-16 · Barack Wamkaya Wanjawa, Lawrence Muchemi, Evans Miriti

Processing low-resource languages, such as Kiswahili, using machine learning is difficult due to lack of adequate training data. However, such low-resource languages are still important for human communication and are al…

Question Answering

Data centric approach to Chinese Medical Speech Recognition

2021-10-01 · ROCLING 2021 10 · Sheng-Luen Chung, Yi-Shiuan Li, Hsien-Wei Ting

Concerning the development of Chinese medical speech recognition technology, this study re-addresses earlier encountered issues in accordance with the process of Machine Learning Engineering for Production (MLOps) from a…

Data Augmentationspeech-recognitionSpeech Recognition

Huqariq: A Multilingual Speech Corpus of Native Languages of Peru forSpeech Recognition

2022-06-01 · LREC 2022 6 · Rodolfo Zevallos, Luis Camacho, Nelsi Melgarejo

The Huqariq corpus is a multilingual collection of speech from native Peruvian languages. The transcribed corpus is intended for the research and development of speech technologies to preserve endangered languages in Per…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3

Huqariq: A Multilingual Speech Corpus of Native Languages of Peru for Speech Recognition

2022-07-12 · Rodolfo Zevallos, Luis Camacho, Nelsi Melgarejo

The Huqariq corpus is a multilingual collection of speech from native Peruvian languages. The transcribed corpus is intended for the research and development of speech technologies to preserve endangered languages in Per…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3