End-to-End Automatic Speech Recognition Integrated With CTC-Based Voice Activity Detection
This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification (CTC) and its extension of CTC/attention architectures. As opposed to an attention-based architecture, input-synchronous label prediction can be performed based on a greedy search with the CTC (pre-)softmax output. This prediction includes consecutive long blank labels, which can be regarded as a non-speech region. We use the labels as a cue for detecting speech segments with simple thresholding. The threshold value is directly related to the length of a non-speech region, which is more intuitive and easier to control than conventional VAD hyperparameters. Experimental results on unsegmented data show that the proposed method outperformed the baseline methods using the conventional energy-based and neural-network-based VAD methods and achieved an RTF less than 0.2. The proposed method is publicly available.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations
Although automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper pr…
Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+3Towards Voice Reconstruction from EEG during Imagined Speech
Translating imagined speech from human brain activity into voice is a challenging and absorbing research issue that can provide new means of human communication via brain signals. Endeavors toward reconstructing speech f…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderEEG+4Tiny Noise-Robust Voice Activity Detector for Voice Assistants
Voice Activity Detection (VAD) in the presence of background noise remains a challenging problem in speech processing. Accurate VAD is essential in automatic speech recognition, voice-to-text, conversational agents, etc,…
Speech RecognitionActivity DetectionX-Vector based voice activity detection for multi-genre broadcast speech-to-text
Voice Activity Detection (VAD) is a fundamental preprocessing step in automatic speech recognition. This is especially true within the broadcast industry where a wide variety of audio materials and recording conditions a…
Action DetectionActivity DetectionAudio ClassificationAutomatic Speech Recognition+4STC Speaker Recognition Systems for the VOiCES From a Distance Challenge
This paper presents the Speech Technology Center (STC) speaker recognition (SR) systems submitted to the VOiCES From a Distance challenge 2019. The challenge's SR task is focused on the problem of speaker recognition in …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMetric Learning+4