VSANet: Real-time Speech Enhancement Based on Voice Activity Detection and Causal Spatial Attention
The deep learning-based speech enhancement (SE) methods always take the clean speech's waveform or time-frequency spectrum feature as the learning target, and train the deep neural network (DNN) by reducing the error loss between the DNN's output and the target. This is a conventional single-task learning paradigm, which has been proven to be effective, but we find that the multi-task learning framework can improve SE performance. Specifically, we design a framework containing a SE module and a voice activity detection (VAD) module, both of which share the same encoder, and the whole network is optimized by the weighted loss of the two modules. Moreover, we design a causal spatial attention (CSA) block to promote the representation capability of DNN. Combining the VAD aided multi-task learning framework and CSA block, our SE network is named VSANet. The experimental results prove the benefits of multi-task learning and the CSA block, which give VSANet an excellent SE performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionMulti-Task LearningSpeech EnhancementSimilar Papers 제목 키워드 기반
Speech enhancement aided end-to-end multi-task learning for voice activity detection
Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address…
Action DetectionActivity DetectionDecoderMulti-Task Learning+2VoiceID Loss: Speech Enhancement for Speaker Verification
In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancem…
Speaker VerificationSpeech EnhancementReal-Time System for Audio-Visual Target Speech Enhancement
We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…
Audio-Visual Speech RecognitionSpeech EnhancementImproved Speech Enhancement with the Wave-U-Net
We study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment. This end-to-end learning method for audio source separatio…
Audio Source SeparationSpeech Enhancementspeech-recognitionSpeech RecognitionImproved Speech Enhancement with the Wave-U-Net
We study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment. This end-to-end learning method for audio source separati…
Audio Source SeparationSpeech Enhancementspeech-recognitionSpeech Recognition