paper-with-me

Papers

A Deep Representation Learning-based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder

2023-12-15 · Yang Xiang, Jingguang Tian, Xinhui Hu, Xinkang Xu, ZhaoHui Yin

Generally, the performance of deep neural networks (DNNs) heavily depends on the quality of data representation learning. Our preliminary work has emphasized the significance of deep representation learning (DRL) in the context of speech enhancement (SE) applications. Specifically, our initial SE algorithm employed a gated recurrent unit variational autoencoder (VAE) with a Gaussian distribution to enhance the performance of certain existing SE systems. Building upon our preliminary framework, this paper introduces a novel approach for SE using deep complex convolutional recurrent networks with a VAE (DCCRN-VAE). DCCRN-VAE assumes that the latent variables of signals follow complex Gaussian distributions that are modeled by DCCRN, as these distributions can better capture the behaviors of complex signals. Additionally, we propose the application of a residual loss in DCCRN-VAE to further improve the quality of the enhanced speech. {Compared to our preliminary work, DCCRN-VAE introduces a more sophisticated DCCRN structure and probability distribution for DRL. Furthermore, in comparison to DCCRN, DCCRN-VAE employs a more advanced DRL strategy. The experimental results demonstrate that the proposed SE algorithm outperforms both our preliminary SE framework and the state-of-the-art DCCRN SE method in terms of scale-invariant signal-to-distortion ratio, speech quality, and speech intelligibility.

📄 PDF Abstract BibTeX arXiv:2312.09620

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeech Enhancement

Similar Papers 제목 키워드 기반

Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses

2021-02-03 · Shengkui Zhao, Trung Hieu Nguyen, Bin Ma

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures wi…

DecoderSpeech DenoisingSpeech Enhancement

Single Channel Speech Enhancement Using Temporal Convolutional Recurrent Neural Networks

2020-02-02 · Jingdong Li, HUI ZHANG, Xueliang Zhang, Changliang Li

In recent decades, neural network based methods have significantly improved the performace of speech enhancement. Most of them estimate time-frequency (T-F) representation of target speech directly or indirectly, then re…

Speech Enhancement

Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network for Speech Enhancement

2021-04-12 · Liming Zhou, Yongyu Gao, Ziluo Wang, Jiwei Li 외

Speech enhancement has benefited from the success of deep learning in terms of intelligibility and perceptual quality. Conventional time-frequency (TF) domain methods focus on predicting TF-masks or speech spectrum,via a…

DecoderSpeech Enhancement

DCCRGAN: Deep Complex Convolution Recurrent Generator Adversarial Network for Speech Enhancement

2020-12-19 · Huixiang Huang, Renjie Wu, Jingbiao Huang, Jucai Lin 외

Generative adversarial network (GAN) still exists some problems in dealing with speech enhancement (SE) task. Some GAN-based systems adopt the same structure from Pixel-to-Pixel directly without special optimization. The…

Generative Adversarial NetworkSpeech Enhancement

Constrained Convolutional-Recurrent Networks to Improve Speech Quality with Low Impact on Recognition Accuracy

2018-02-16 · Rasool Fakoor, Xiaodong He, Ivan Tashev, Shuayb Zarar

For a speech-enhancement algorithm, it is highly desirable to simultaneously improve perceptual quality and recognition rate. Thanks to computational costs and model complexities, it is challenging to train a model that …

DecoderLanguage ModelingLanguage ModellingSpeech Enhancement