paper-with-me

Speech Recognition

65개 벤치마크 · 논문 7,125편 · 이 태스크의 논문 보기 →

Benchmarks

LibriSpeech test-clean

결과 66개

LibriSpeech test-other

결과 53개

Switchboard + Hub500

결과 30개

TIMIT

결과 22개

AISHELL-1

결과 20개

WSJ eval92

결과 17개

Common Voice German

결과 14개

TUDA

결과 9개

Common Voice French

결과 8개

Common Voice Spanish

결과 8개

MediaSpeech

결과 8개

SLUE

결과 8개

VietMed

결과 8개

WenetSpeech

결과 8개

Common Voice

결과 5개

EasyCom

결과 5개

GigaSpeech DEV

결과 5개

GigaSpeech TEST

결과 5개

Hub5'00 SwitchBoard

결과 5개

CHiME-6 dev_gss12

결과 4개

LRS3-TED

결과 4개

Tedlium

결과 4개

WSJ dev93

결과 4개

CHiME-6 eval

결과 3개

Common Voice vi

결과 3개

Fongbe audio

결과 3개

SPGISpeech

결과 3개

Speech Commands

결과 3개

VIVOS

결과 3개

WSJ eval93

결과 3개

AISHELL-2

결과 2개

AMI IMH

결과 2개

AMI SDM1

결과 2개

Common Voice English

결과 2개

Common Voice Italian

결과 2개

LibriCSS

결과 2개

TED-LIUM

결과 2개

AISHELL-2 Test IOS

결과 1개

AISHELL-2 Test Mic

결과 1개

CALLHOME En

결과 1개

CAS-VSR-S101

결과 1개

Common Voice Frisian

결과 1개

Common Voice Japanese

결과 1개

Common Voice Russian

결과 1개

GigaSpeech

결과 1개

Hub5'00 CallHome

결과 1개

Hub5'00 FISHER-SWBD

결과 1개

LRS2

결과 1개

Switchboard (300hr)

결과 1개

Switchboard CallHome

결과 1개

Switchboard SWBD

결과 1개

Most implemented

Listen, Attend and Spell

2015-08-05 · 구현 40개

Papers

Do speech foundation models really learn words?

2026-09-09 · Robin Huo, Ewan Dunbar arxiv

Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understa…

Speech Recognition

Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services

2026-09-09 · Yonghyun Jun, Jimin Lee, Hwan Chang, Dongho Shin 외 arxiv

Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments su…

Speech Recognition

$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants

2026-09-09 · Heyang Liu, Jiayi Huang, Wenyang Xiao, Ziyang Cheng 외 arxiv

The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable pe…

Speech Recognition

StreamAlign: Streaming Text-Aligned Speech Tokenization

2026-09-09 · Kang-wook Kim, Jinyoung Park, Jinsoo Kim, Sehun Lee 외 arxiv

Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognitio…

Speech Recognition

SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

2026-09-09 · Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He 외 arxiv

The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (…

Emotion RecognitionSpeaker RecognitionSpeech RecognitionQuestion Answering

Qwen-Audio-3.0-ASR Technical Report

2026-09-07 · Chuanmeng Bian, Daren Chen, Peixin Chen, Zhigao Chen 외 arxiv

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). …

Speech Recognition

전체 7,125편 보기 →