VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition
While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed methods, (partial) overlapping inference are shown to be effective on long-form decoding. For both methods, word error rate (WER) decreases monotonically when overlapping percentage decreases. Setting aside computational cost, the setup with 50% overlapping during inference can achieve the best performance. However, a lower overlapping percentage has an advantage of fast inference speed. In this paper, we first conduct comprehensive experiments comparing overlapping inference and partial overlapping inference with various configurations. We then propose Voice-Activity-Detection Overlapping Inference to provide a trade-off between WER and computation cost. Results show that the proposed method can achieve a 20% relative computation cost reduction on Librispeech and Microsoft Speech Language Translation long-form corpus while maintaining the WER performance when comparing to the best performing overlapping inference algorithm. We also propose Soft-Match to compensate for similar words mis-aligned problem.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Formspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
An End-to-End Architecture for Keyword Spotting and Voice Activity Detection
We propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection. We develop novel inference algorithms for an end-to-end Recurrent Neural Network trained with the Conn…
Action DetectionActivity DetectionGeneral ClassificationKeyword SpottingA Convolutional Neural Network Smartphone App for Real-Time Voice Activity Detection
This paper presents a smartphone app that performs real-time voice activity detection based on convolutional neural network. Real-time implementation issues are discussed showing how the slow inference time associated wi…
Action DetectionActivity DetectionAudio Signal RecognitionNoise EstimationMultitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0
Self-supervised learning approaches have lately achieved great success on a broad spectrum of machine learning problems. In the field of speech processing, one of the most successful recent self-supervised models is wav2…
Action DetectionActivity DetectionChange DetectionSelf-Supervised LearningSpeaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
In speaker diarization, traditional clustering-based methods remain widely used in real-world applications. However, these methods struggle with the complex distribution of speaker embeddings and overlapping speech segme…
Action DetectionActivity DetectionClusteringCommunity Detection+3Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario
Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle o…
Action DetectionActivity DetectionBinary ClassificationClustering+2