paper-with-me

홈 › Papers

Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

2024-06-14 · Haoyu Wang, Guoqiang Hu, Guodong Lin, Wei-Qiang Zhang, Jian Li

As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its encoder-decoder structure hinders its application to streaming speech recognition. In this paper, we introduce Simul-Whisper, which uses the time alignment embedded in Whisper's cross-attention to guide auto-regressive decoding and achieve chunk-based streaming ASR without any fine-tuning of the pre-trained model. Furthermore, we observe the negative effect of the truncated words at the chunk boundaries on the decoding results and propose an integrate-and-fire-based truncation detection model to address this issue. Experiments on multiple languages and Whisper architectures show that Simul-Whisper achieves an average absolute word error rate degradation of only 1.46% at a chunk size of 1 second, which significantly outperforms the current state-of-the-art baseline.

📄 PDF Abstract BibTeX arXiv:2406.10052

Code (1)

backspacetg/simul_whisper 공식 구현 pytorch

Tasks

Decoderspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Turning Whisper into Real-Time Transcription System

2023-07-27 · Dominik Macháček, Raj Dabre, Ondřej Bojar

Whisper is one of the recent state-of-the-art multilingual speech recognition and translation models, however, it is not designed for real time transcription. In this paper, we build on top of Whisper and create Whisper-…

speech-recognitionSpeech RecognitionTranslation

WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

2026-04-28 · Erfan Ramezani, Mohammad Mahdi Giahi, Mohammad Erfan Zarabadipour, Amir Reza Yosefian 외 arxiv

Real-time automatic speech recognition (ASR) systems face a fundamental trade-off between transcription accuracy and computational efficiency, particularly when deploying large-scale transformer models like Whisper. Exis…

Computational EfficiencySpeech RecognitionActivity Detection

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition

2025-06-04 · Yinfeng Xia, Huiyan Li, Chenyang Le, Manhong Wang 외

Applying large pre-trained speech models like Whisper has shown promise in reducing training costs for various speech tasks. However, integrating these models into streaming systems remains a challenge. This paper presen…

speech-recognitionSpeech Recognition

Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion

2026-02-28 · Sen Zhang, Jianguo Wei, Wenhuan Lu, Xianghu Yue 외 arxiv

The Transformer-based Whisper model has achieved state-of-the-art performance in Automatic Speech Recognition (ASR). However, its Multi-Head Attention (MHA) mechanism results in significant GPU memory consumption due to …

Speech Recognition

WhisperRT -- Turning Whisper into a Causal Streaming Model

2025-08-17 · Tomer Krichli, Bhiksha Raj, Joseph Keshet arxiv

Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not des…

Speech Recognition