paper-with-me

홈 › Papers

CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions

2024-08-29 · Laurin Wagner, Bernhard Thallinger, Mario Zusag

We demonstrate that carefully adjusting the tokenizer of the Whisper speech recognition model significantly improves the precision of word-level timestamps when applying dynamic time warping to the decoder's cross-attention scores. We fine-tune the model to produce more verbatim speech transcriptions and employ several techniques to increase robustness against multiple speakers and background noise. These adjustments achieve state-of-the-art performance on benchmarks for verbatim speech transcription, word segmentation, and the timed detection of filler events, and can further mitigate transcription hallucinations. The code is available open https://github.com/nyrahealth/CrisperWhisper.

📄 PDF Abstract BibTeX arXiv:2408.16589

Code (1)

nyrahealth/crisperwhisper 공식 구현 pytorch

Tasks

Dynamic Time Warpingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

2026-07-21 · Laurin Wagner, Mario Zusag, Bernhard Thallinger hf

Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60%…

Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection

2025-05-29 · Griffin Dietz Smith, Dianna Yee, Jennifer King Chen, Leah Findlater

Identifying mistakes (i.e., miscues) made while reading aloud is commonly approached post-hoc by comparing automatic speech recognition (ASR) transcriptions to the target reading text. However, post-hoc methods perform p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Learning to Jointly Transcribe and Subtitle for End-to-End Spontaneous Speech Recognition

2022-10-14 · Jakob Poncelet, Hugo Van hamme

TV subtitles are a rich source of transcriptions of many types of speech, ranging from read speech in news reports to conversational and spontaneous speech in talk shows and soaps. However, subtitles are not verbatim (i.…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1

Automatic Detection of Everyday Social Behaviours and Environments from Verbatim Transcripts of Daily Conversations

2019-07-22 · Kristina Y. Yordanova, Burcu Demiray, Matthias R. Mehl, Mike Martin

Coding in social sciences is a process that involves the categorisation of qualitative or quantitative data in order to facilitate further analysis. Coding is usually a manual process that involves a lot of effort and ti…

Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling

2025-02-05 · Jakob Poncelet, Hugo Van hamme

The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for low-resource languages and dialects. Thi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition