Automatic Speech Recognition (ASR)
9개 벤치마크 · 논문 3,012편 · 이 태스크의 논문 보기 →
Benchmarks
LRS2
LRS3-TED
RealMAN
Sagalee
HUI speech corpus
M-AILabs speech dataset
VoxPopuli
Voxforge German
Most implemented
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Conformer: Convolution-augmented Transformer for Speech Recognition
Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces
Papers
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion ClassificationExpressive Speech Synthesis+5WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionLightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
Overlapping speech remains a major challenge for automatic speech recognition (ASR) in real-world applications, particularly in broadcast media with dynamic, multi-speaker interactions. We propose a light-weight, target-…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionEnd-to-End Spoken Grammatical Error Correction
Grammatical Error Correction (GEC) and feedback play a vital role in supporting second language (L2) learners, educators, and examiners. While written GEC is well-established, spoken GEC (SGEC), aiming to provide feedbac…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Grammatical Error Correctionspeech-recognition+1AI-Generated Song Detection via Lyrics Transcripts
The recent rise in capabilities of AI-based music generation tools has created an upheaval in the music industry, necessitating the creation of accurate methods to detect such AI-generated content. This can be done using…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Music Generationspeech-recognition+1Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
Automatic Speech Recognition (ASR) has reached impressive accuracy for high-resource languages, yet its utility in linguistic fieldwork remains limited. Recordings collected in fieldwork contexts present unique challenge…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition