paper-with-me

홈 › Papers

DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting

2023-05-21 · Shubo Lv, Xiong Wang, Sining Sun, Long Ma, Lei Xie

Real-world complex acoustic environments especially the ones with a low signal-to-noise ratio (SNR) will bring tremendous challenges to a keyword spotting (KWS) system. Inspired by the recent advances of neural speech enhancement and context bias in speech recognition, we propose a robust audio context bias based DCCRN-KWS model to address this challenge. We form the whole architecture as a multi-task learning framework for both denosing and keyword spotting, where the DCCRN encoder is connected with the KWS model. Helped with the denoising task, we further introduce an audio context bias module to leverage the real keyword samples and bias the network to better iscriminate keywords in noisy conditions. Feature merge and complex context linear modules are also introduced to strength such discrimination and to effectively leverage contextual information respectively. Experiments on the internal challenging dataset and the HIMIYA public dataset show that our DCCRN-KWS system is superior in performance, while ablation study demonstrates the good design of the whole model.

📄 PDF Abstract BibTeX arXiv:2305.12331

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingKeyword SpottingMulti-Task LearningSmall-Footprint Keyword SpottingSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

S-DCCRN: Super Wide Band DCCRN with learnable complex feature for speech enhancement

2021-11-16 · Shubo Lv, Yihui Fu, Mengtao Xing, Jiayao Sun 외

In speech enhancement, complex neural network has shown promising performance due to their effectiveness in processing complex-valued spectrum. Most of the recent speech enhancement approaches mainly focus on wide-band s…

16kDenoisingSpeech DenoisingSpeech Enhancement

DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement

2021-06-16 · Shubo Lv, Yanxin Hu, Shimin Zhang, Lei Xie

Deep complex convolution recurrent network (DCCRN), which extends CRN with complex structure, has achieved superior performance in MOS evaluation in Interspeech 2020 deep noise suppression challenge (DNS2020). This paper…

DecoderSpeech Enhancement

spatial-dccrn: dccrn equipped with frame-level angle feature and hybrid filtering for multi-channel speech enhancement

2022-10-17 · Shubo Lv, Yihui Fu, Yukai Jv, Lei Xie 외

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network ba…

DenoisingSpeech Enhancement

A Deep Representation Learning-based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder

2023-12-15 · Yang Xiang, Jingguang Tian, Xinhui Hu, Xinkang Xu 외

Generally, the performance of deep neural networks (DNNs) heavily depends on the quality of data representation learning. Our preliminary work has emphasized the significance of deep representation learning (DRL) in the …

Representation LearningSpeech Enhancement

Causal Signal-Based DCCRN with Overlapped-Frame Prediction for Online Speech Enhancement

2023-09-07 · Julitta Bartolewska, Stanisław Kacprzak, Konrad Kowalczyk

The aim of speech enhancement is to improve speech signal quality and intelligibility from a noisy microphone signal. In many applications, it is crucial to enable processing with small computational complexity and minim…

Speech Enhancement