ASR Benchmarking: Need for a More Representative Conversational Dataset
Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational environments, where speech is often unstructured and contains disfluencies such as pauses, interruptions, and diverse accents. In this study, we introduce a multilingual conversational dataset, derived from TalkBank, consisting of unstructured phone conversation between adults. Our results show a significant performance drop across various state-of-the-art ASR models when tested in conversational settings. Furthermore, we observe a correlation between Word Error Rate and the presence of speech disfluencies, highlighting the critical need for more realistic, conversational ASR benchmarks.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarkingspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
WER We Stand: Benchmarking Urdu ASR Models
This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarkingspeech-recognition+2GiCCS: A German in-Context Conversational Similarity Benchmark
The Semantic textual similarity (STS) task is commonly used to evaluate the semantic representations that language models (LMs) learn from texts, under the assumption that good-quality representations will yield accurate…
BenchmarkingSemantic Textual SimilaritySTSA Case for Dataset Specific Profiling
Data-driven science is an emerging paradigm where scientific discoveries depend on the execution of computational AI models against rich, discipline-specific datasets. With modern machine learning frameworks, anyone can …
BenchmarkingModel SelectionEmotion Recognition in Sign Language Conversation
Emotion Recognition in Conversation is a core component of affective computing, while current resources of sign language emotion datasets primarily focus on isolated sentences and lack conversational context. Models trai…
Emotion Recognition in ConversationAI applications in forest monitoring need remote sensing benchmark datasets
With the rise in high resolution remote sensing technologies there has been an explosion in the amount of data available for forest monitoring, and an accompanying growth in artificial intelligence applications to automa…
Benchmarking