On training targets for noise-robust voice activity detection
The task of voice activity detection (VAD) is an often required module in various speech processing, analysis and classification tasks. While state-of-the-art neural network based VADs can achieve great results, they often exceed computational budgets and real-time operating requirements. In this work, we propose a computationally efficient real-time VAD network that achieves state-of-the-art results on several public real recording datasets. We investigate different training targets for the VAD and show that using the segmental voice-to-noise ratio (VNR) is a better and more noise-robust training target than the clean speech level based VAD. We also show that multi-target training improves the performance further.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionSimilar Papers 제목 키워드 기반
A Convolutional Neural Network Smartphone App for Real-Time Voice Activity Detection
This paper presents a smartphone app that performs real-time voice activity detection based on convolutional neural network. Real-time implementation issues are discussed showing how the slow inference time associated wi…
Action DetectionActivity DetectionAudio Signal RecognitionNoise EstimationTiny Noise-Robust Voice Activity Detector for Voice Assistants
Voice Activity Detection (VAD) in the presence of background noise remains a challenging problem in speech processing. Accurate VAD is essential in automatic speech recognition, voice-to-text, conversational agents, etc,…
Speech RecognitionActivity DetectionAdversarial Multi-Task Deep Learning for Noise-Robust Voice Activity Detection with Low Algorithmic Delay
Voice Activity Detection (VAD) is an important pre-processing step in a wide variety of speech processing systems. VAD should in a practical application be able to detect speech in both noisy and noise-free environments,…
Action DetectionActivity DetectionMulti-Task LearningrVAD: An Unsupervised Segment-Based Robust Voice Activity Detection Method
This paper presents an unsupervised segment-based method for robust voice activity detection (rVAD). The method consists of two passes of denoising followed by a voice activity detection (VAD) stage. In the first pass, h…
Action DetectionActivity DetectionDenoisingSpeaker Verification+1An End-to-End Architecture for Keyword Spotting and Voice Activity Detection
We propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection. We develop novel inference algorithms for an end-to-end Recurrent Neural Network trained with the Conn…
Action DetectionActivity DetectionGeneral ClassificationKeyword Spotting