Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model
With fewer feature dimensions, filter banks are often used in light-weight full-band speech enhancement models. In order to further enhance the coarse speech in the sub-band domain, it is necessary to apply a post-filtering for harmonic retrieval. The signal processing-based comb filters used in RNNoise and PercepNet have limited performance and may cause speech quality degradation due to inaccurate fundamental frequency estimation. To tackle this problem, we propose a learnable comb filter to enhance harmonics. Based on the sub-band model, we design a DNN-based fundamental frequency estimator to estimate the discrete fundamental frequencies and a comb filter for harmonic enhancement, which are trained via an end-to-end pattern. The experiments show the advantages of our proposed method over PecepNet and DeepFilterNet.
Code (0)
등록된 구현이 없습니다.
Tasks
RetrievalSpeech EnhancementSimilar Papers 제목 키워드 기반
HGCN: Harmonic gated compensation network for speech enhancement
Mask processing in the time-frequency (T-F) domain through the neural network has been one of the mainstreams for single-channel speech enhancement. However, it is hard for most models to handle the situation when harmon…
Action DetectionActivity DetectionSpeech EnhancementHarmonic Detection from Noisy Speech with Auditory Frame Gain for Intelligibility Enhancement
This paper introduces a novel (HDAG - Harmonic Detection for Auditory Gain) method for speech intelligibility enhancement in noisy scenarios. In the proposed scheme, a series of selective Gammachirp filters are adopted t…
Improving Low-Light Image Recognition Performance Based on Image-adaptive Learnable Module
In recent years, significant progress has been made in image recognition technology based on deep neural networks. However, improving recognition performance under low-light conditions remains a significant challenge. Th…
DeepFilterNet2: Towards Real-Time Speech Enhancement on Embedded Devices for Full-Band Audio
Deep learning-based speech enhancement has seen huge improvements and recently also expanded to full band audio (48 kHz). However, many approaches have a rather high computational complexity and require big temporal buff…
CPUData AugmentationSpeech EnhancementHarmonic Convolutional Networks based on Discrete Cosine Transform
Convolutional neural networks (CNNs) learn filters in order to capture local correlation patterns in feature space. We propose to learn these filters as combinations of preset spectral filters defined by the Discrete Cos…
Edge Detectionimage-classificationImage Classificationobject-detection+2