paper-with-me

Papers

VAD-free Streaming Hybrid CTC/Attention ASR for Unsegmented Recording

2021-07-15 · Hirofumi Inaguma, Tatsuya Kawahara

In this work, we propose novel decoding algorithms to enable streaming automatic speech recognition (ASR) on unsegmented long-form recordings without voice activity detection (VAD), based on monotonic chunkwise attention (MoChA) with an auxiliary connectionist temporal classification (CTC) objective. We propose a block-synchronous beam search decoding to take advantage of efficient batched output-synchronous and low-latency input-synchronous searches. We also propose a VAD-free inference algorithm that leverages CTC probabilities to determine a suitable timing to reset the model states to tackle the vulnerability to long-form data. Experimental evaluations demonstrate that the block-synchronous decoding achieves comparable accuracy to the label-synchronous one. Moreover, the VAD-free inference can recognize long-form speech robustly for up to a few hours.

📄 PDF Abstract BibTeX arXiv:2107.07509

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Formspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Segmentation-Free Streaming Machine Translation

2023-09-26 · Javier Iranzo-Sánchez, Jorge Iranzo-Sánchez, Adrià Giménez, Jorge Civera 외

Streaming Machine Translation (MT) is the task of translating an unbounded input text stream in real-time. The traditional cascade approach, which combines an Automatic Speech Recognition (ASR) and an MT system, relies o…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSegmentation+4

CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

2020-04-20 · Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent 외

Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits the previous CHiME-5 challenge and furt…

speaker-diarizationSpeaker DiarizationSpeech Enhancementspeech-recognition+2

SURT 2.0: Advances in Transducer-based Multi-talker Speech Recognition

2023-06-18 · Desh Raj, Daniel Povey, Sanjeev Khudanpur

The Streaming Unmixing and Recognition Transducer (SURT) model was proposed recently as an end-to-end approach for continuous, streaming, multi-talker speech recognition (ASR). Despite impressive results on multi-turn me…

DecoderDomain Adaptationspeech-recognitionSpeech Recognition

Turning Whisper into Real-Time Transcription System

2023-07-27 · Dominik Macháček, Raj Dabre, Ondřej Bojar

Whisper is one of the recent state-of-the-art multilingual speech recognition and translation models, however, it is not designed for real time transcription. In this paper, we build on top of Whisper and create Whisper-…

speech-recognitionSpeech RecognitionTranslation

ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data

2025-09-08 · Vladislav Stankov, Matyáš Kopp, Ondřej Bojar arxiv

We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the sound recordings of the Czech parliamen…

Speech RecognitionSpeech Synthesis