paper-with-me

Papers

Speech enhancement aided end-to-end multi-task learning for voice activity detection

2020-10-23 · Xu Tan, Xiao-Lei Zhang

Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address this issue, here we propose a speech enhancement aided end-to-end multi-task model for VAD. The model has two decoders, one for speech enhancement and the other for VAD. The two decoders share the same encoder and speech separation network. Unlike the direct thought that takes two separated objectives for VAD and speech enhancement respectively, here we propose a new joint optimization objective -- VAD-masked scale-invariant source-to-distortion ratio (mSI-SDR). mSI-SDR uses VAD information to mask the output of the speech enhancement decoder in the training process. It makes the VAD and speech enhancement tasks jointly optimized not only at the shared encoder and separation network, but also at the objective level. It also satisfies real-time working requirement theoretically. Experimental results show that the multi-task method significantly outperforms its single-task VAD counterpart. Moreover, mSI-SDR outperforms SI-SDR in the same multi-task setting.

📄 PDF Abstract BibTeX arXiv:2010.12484

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionDecoderMulti-Task LearningSpeech EnhancementSpeech Separation

Similar Papers 제목 키워드 기반

VSANet: Real-time Speech Enhancement Based on Voice Activity Detection and Causal Spatial Attention

2023-10-11 · Yuewei Zhang, Huanbin Zou, Jie Zhu

The deep learning-based speech enhancement (SE) methods always take the clean speech's waveform or time-frequency spectrum feature as the learning target, and train the deep neural network (DNN) by reducing the error los…

Action DetectionActivity DetectionMulti-Task LearningSpeech Enhancement

VoiceID Loss: Speech Enhancement for Speaker Verification

2019-04-07 · Suwon Shon, Hao Tang, James Glass

In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancem…

Speaker VerificationSpeech Enhancement

Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation

2025-12-07 · Jisoo Park, Seonghak Lee, Guisik Kim, Taewoo Kim 외 arxiv

Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivati…

Speech EnhancementSpeech ExtractionSpeech Separation

Improved Speech Enhancement with the Wave-U-Net

2018-11-27 · Craig Macartney, Tillman Weyde

We study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment. This end-to-end learning method for audio source separatio…

Audio Source SeparationSpeech Enhancementspeech-recognitionSpeech Recognition

Improved Speech Enhancement with the Wave-U-Net

2018-10-22 · Anonymous

We study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment. This end-to-end learning method for audio source separati…

Audio Source SeparationSpeech Enhancementspeech-recognitionSpeech Recognition