paper-with-me

홈 › Papers

Earnings-22: A Practical Benchmark for Accents in the Wild

2022-03-29 · Miguel Del Rio, Peter Ha, Quinten McNamara, Corey Miller, Shipra Chandra

Modern automatic speech recognition (ASR) systems have achieved superhuman Word Error Rate (WER) on many common corpora despite lacking adequate performance on speech in the wild. Beyond that, there is a lack of real-world, accented corpora to properly benchmark academic and commercial models. To ensure this type of speech is represented in ASR benchmarking, we present Earnings-22, a 125 file, 119 hour corpus of English-language earnings calls gathered from global companies. We run a comparison across 4 commercial models showing the variation in performance when taking country of origin into consideration. Looking at hypothesis transcriptions, we explore errors common to all ASR systems tested. By examining Individual Word Error Rate (IWER), we find that key speech features impact model performance more for certain accents than others. Earnings-22 provides a free-to-use benchmark of real-world, accented audio to bridge academic and industrial research.

📄 PDF Abstract BibTeX arXiv:2203.15591

Code (1)

revdotcom/speech-datasets 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarkingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Earnings-21: A Practical Benchmark for ASR in the Wild

2021-04-22 · Miguel Del Rio, Natalie Delworth, Ryan Westerman, Michelle Huang 외

Commonly used speech corpora inadequately challenge academic and commercial ASR systems. In particular, speech corpora lack metadata needed for detailed analysis and WER measurement. In response, we present Earnings-21, …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition

2026-05-05 · Busayo Awobade, Gabrial Zencha Ashungafac, Tobi Olatunji arxiv

Recent large language models (LLMs) show strong speech recognition and translation capabilities for high-resource languages. However, African languages remain dramatically underrepresented in benchmarks, limiting their p…

Speech Recognition

Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild

2026-03-28 · Berkin Durmus, Chen Cen, Eduardo Pacheco, Arda Okan 외 arxiv

The accuracy frontier of speech-to-text systems has plateaued on academic benchmarks.1 In contrast, industrial benchmarks and adoption in high-stakes domains suggest otherwise. We hypothesize that the primary difference …

Speech Recognition

Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units

2026-08-27 · Robin San Roman, Manel Khentout, Tu Anh Nguyen, Paul Michel 외 arxiv

Representation learning has attracted great atten- tion and managed to reach good performances as a pretraining method for downstream tasks or as a first step towards unsu- pervised speech modeling. Yet, little is known …

Representation Learning

AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents

2024-02-02 · Abraham Toluwase Owodunni, Aditya Yadavalli, Chris Chinenye Emezue, Tobi Olatunji 외

Despite advancements in speech recognition, accented speech remains challenging. While previous approaches have focused on modeling techniques or creating accented speech datasets, gathering sufficient data for the multi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1