paper-with-me

Papers

Frequency Gating: Improved Convolutional Neural Networks for Speech Enhancement in the Time-Frequency Domain

2020-11-08 · Koen Oostermeijer, Qing Wang, Jun Du

One of the strengths of traditional convolutional neural networks (CNNs) is their inherent translational invariance. However, for the task of speech enhancement in the time-frequency domain, this property cannot be fully exploited due to a lack of invariance in the frequency direction. In this paper we propose to remedy this inefficiency by introducing a method, which we call Frequency Gating, to compute multiplicative weights for the kernels of the CNN in order to make them frequency dependent. Several mechanisms are explored: temporal gating, in which weights are dependent on prior time frames, local gating, whose weights are generated based on a single time frame and the ones adjacent to it, and frequency-wise gating, where each kernel is assigned a weight independent of the input data. Experiments with an autoencoder neural network with skip connections show that both local and frequency-wise gating outperform the baseline and are therefore viable ways to improve CNN-based speech enhancement neural networks. In addition, a loss function based on the extended short-time objective intelligibility score (ESTOI) is introduced, which we show to outperform the standard mean squared error (MSE) loss function.

📄 PDF Abstract BibTeX arXiv:2011.04092

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques

2024-02-22 · Changjiang Zhao, Shulin He, Xueliang Zhang

Speech enhancement aims to improve speech quality and intelligibility, especially in noisy environments where background noise degrades speech signals. Currently, deep learning methods achieve great success in speech enh…

Speech Enhancement

Single Channel Speech Enhancement Using Temporal Convolutional Recurrent Neural Networks

2020-02-02 · Jingdong Li, HUI ZHANG, Xueliang Zhang, Changliang Li

In recent decades, neural network based methods have significantly improved the performace of speech enhancement. Most of them estimate time-frequency (T-F) representation of target speech directly or indirectly, then re…

Speech Enhancement

A Dual-Staged Context Aggregation Method Towards Efficient End-To-End Speech Enhancement

2019-08-18 · Kai Zhen, Mi Suk Lee, Minje Kim

In speech enhancement, an end-to-end deep neural network converts a noisy speech signal to a clean speech directly in time domain without time-frequency transformation or mask estimation. However, aggregating contextual …

Speech Enhancement

Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses

2021-02-03 · Shengkui Zhao, Trung Hieu Nguyen, Bin Ma

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures wi…

DecoderSpeech DenoisingSpeech Enhancement

FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network

2022-03-15 · Zehua Zhang, Lu Zhang, Xuyi Zhuang, Yukun Qian 외

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time en…

DenoisingSpeech DenoisingSpeech Enhancement