paper-with-me

홈 › Papers

AfriSpeech-MultiBench: A Verticalized Multidomain Multicountry Benchmark Suite for African Accented English ASR

2025-11-18 · Gabrial Zencha Ashungafac, Mardhiyah Sanni, Busayo Awobade, Alex Gichamba, Tobi Olatunji arxiv

Recent advances in speech-enabled AI, including Google's NotebookLM and OpenAI's speech-to-speech API, are driving widespread interest in voice interfaces globally. Despite this momentum, there exists no publicly available application-specific model evaluation that caters to Africa's linguistic diversity. We present AfriSpeech-MultiBench, the first domain-specific evaluation suite for over 100 African English accents across 10+ countries and seven application domains: Finance, Legal, Medical, General dialogue, Call Center, Named Entities and Hallucination Robustness. We benchmark a diverse range of open, closed, unimodal ASR and multimodal LLM-based speech recognition systems using both spontaneous and non-spontaneous speech conversation drawn from various open African accented English speech datasets. Our empirical analysis reveals systematic variation: open-source ASR models excels in spontaneous speech contexts but degrades on noisy, non-native dialogue; multimodal LLMs are more accent-robust yet struggle with domain-specific named entities; proprietary models deliver high accuracy on clean speech but vary significantly by country and domain. Models fine-tuned on African English achieve competitive accuracy with lower latency, a practical advantage for deployment, hallucinations still remain a big problem for most SOTA models. By releasing this comprehensive benchmark, we empower practitioners and researchers to select voice technologies suited to African use-cases, fostering inclusive voice applications for underserved communities.

📄 PDF Abstract BibTeX arXiv:2511.14255

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

2021-07-15 · Paul Pu Liang, Yiwei Lyu, Xiang Fan, Zetian Wu 외

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. It is a challenging yet crucial area with numerous real-world applications in multimedia, affective comput…

Representation Learning

AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR

2023-09-30 · Tobi Olatunji, Tejumade Afonja, Aditya Yadavalli, Chris Chinenye Emezue 외

Africa has a very low doctor-to-patient ratio. At very busy clinics, doctors could see 30+ patients per day -- a heavy patient burden compared with developed countries -- but productivity tools such as clinical automatic…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

MultiZoo & MultiBench: A Standardized Toolkit for Multimodal Deep Learning

2023-06-28 · Paul Pu Liang, Yiwei Lyu, Xiang Fan, Arav Agarwal 외

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. In order to accelerate progress towards understudied modalities and tasks while ensuring real-world robust…

Deep LearningMultimodal Deep Learning

Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond

2025-02-06 · Mardhiyah Sanni, Tassallah Abdullahi, Devendra D. Kayande, Emmanuel Ayodele 외

Speech technologies are transforming interactions across various sectors, from healthcare to call centers and robots, yet their performance on African-accented conversations remains underexplored. We introduce Afrispeech…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Conversation Summarizationspeaker-diarization+3

Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd Counting

2024-02-06 · Mingyue Guo, Binghui Chen, Zhaoyi Yan, YaoWei Wang 외

Multidomain crowd counting aims to learn a general model for multiple diverse datasets. However, deep networks prefer modeling distributions of the dominant domains instead of all domains, which is known as domain bias. …

Crowd Counting