Speech Recognition
65개 벤치마크 · 논문 7,125편 · 이 태스크의 논문 보기 →
Benchmarks
LibriSpeech test-clean
LibriSpeech test-other
Switchboard + Hub500
TIMIT
AISHELL-1
WSJ eval92
Common Voice German
swb_hub_500 WER fullSWBCH
TUDA
Common Voice French
Common Voice Spanish
MediaSpeech
SLUE
VietMed
WenetSpeech
Common Voice
EasyCom
GigaSpeech DEV
GigaSpeech TEST
Hub5'00 SwitchBoard
Libri-Light test-clean
Libri-Light test-other
CHiME-6 dev_gss12
LRS3-TED
Tedlium
WSJ dev93
CHiME-6 eval
Common Voice vi
Fongbe audio
SPGISpeech
Speech Commands
VIVOS
WSJ eval93
AISHELL-2
AMI IMH
AMI SDM1
Common Voice English
Common Voice Italian
Europarl-ASR EN MEP-test
LibriCSS
TED-LIUM
AISHELL-2 Test Android
AISHELL-2 Test IOS
AISHELL-2 Test Mic
CALLHOME En
CALLHOME Spanish Speech
CAS-VSR-S101
Common Voice Frisian
Common Voice Japanese
Common Voice Portuguese
Common Voice Russian
GigaSpeech
Hub5'00 CallHome
Hub5'00 FISHER-SWBD
LRS2
Switchboard (300hr)
Switchboard CallHome
Switchboard SWBD
Most implemented
Communication-Efficient Learning of Deep Networks from Decentralized Data
Listen, Attend and Spell
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Papers
Do speech foundation models really learn words?
Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understa…
Speech RecognitionLeveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services
Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments su…
Speech Recognition$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants
The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable pe…
Speech RecognitionStreamAlign: Streaming Text-Aligned Speech Tokenization
Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognitio…
Speech RecognitionSEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (…
Emotion RecognitionSpeaker RecognitionSpeech RecognitionQuestion AnsweringQwen-Audio-3.0-ASR Technical Report
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). …
Speech Recognition