paper-with-me

Automatic Speech Recognition (ASR)

9개 벤치마크 · 논문 3,012편 · 이 태스크의 논문 보기 →

Benchmarks

LRS2

결과 18개

LRS3-TED

결과 4개

RealMAN

결과 4개

Sagalee

결과 4개

HUI speech corpus

결과 2개

VoxPopuli

결과 2개

Voxforge German

결과 2개

Most implemented

Papers

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

2025-07-17 · Maksim Borisov, Egor Spirin, Daria Diatlova

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion ClassificationExpressive Speech Synthesis+5

WhisperKit: On-device Real-time ASR with Billion-Scale Transformers

2025-07-14 · Atila Orhon, Arda Okan, Berkin Durmus, Zach Nagengast 외

Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR

2025-06-25 · Aleš Pražák, Marie Kunešová, Josef Psutka

Overlapping speech remains a major challenge for automatic speech recognition (ASR) in real-world applications, particularly in broadcast media with dynamic, multi-speaker interactions. We propose a light-weight, target-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

End-to-End Spoken Grammatical Error Correction

2025-06-23 · Mengjie Qian, Rao Ma, Stefano Bannò, Mark J. F. Gales 외

Grammatical Error Correction (GEC) and feedback play a vital role in supporting second language (L2) learners, educators, and examiners. While written GEC is well-established, spoken GEC (SGEC), aiming to provide feedbac…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Grammatical Error Correctionspeech-recognition+1

AI-Generated Song Detection via Lyrics Transcripts

2025-06-23 · Markus Frohmann, Elena V. Epure, Gabriel Meseguer-Brocal, Markus Schedl 외

The recent rise in capabilities of AI-based music generation tools has created an upheaval in the music industry, necessitating the creation of accurate methods to detect such AI-generated content. This can be done using…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Music Generationspeech-recognition+1

Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages

2025-06-20 · Siyu Liang, Gina-Anne Levow

Automatic Speech Recognition (ASR) has reached impressive accuracy for high-resource languages, yet its utility in linguistic fieldwork remains limited. Recordings collected in fieldwork contexts present unique challenge…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

전체 3,012편 보기 →