paper-with-me

Papers

A Dual-Staged Context Aggregation Method Towards Efficient End-To-End Speech Enhancement

2019-08-18 · Kai Zhen, Mi Suk Lee, Minje Kim

In speech enhancement, an end-to-end deep neural network converts a noisy speech signal to a clean speech directly in time domain without time-frequency transformation or mask estimation. However, aggregating contextual information from a high-resolution time domain signal with an affordable model complexity still remains challenging. In this paper, we propose a densely connected convolutional and recurrent network (DCCRN), a hybrid architecture, to enable dual-staged temporal context aggregation. With the dense connectivity and cross-component identical shortcut, DCCRN consistently outperforms competing convolutional baselines with an average STOI improvement of 0.23 and PESQ of 1.38 at three SNR levels. The proposed method is computationally efficient with only 1.38 million parameters. The generalizability performance on the unseen noise types is still decent considering its low complexity, although it is relatively weaker comparing to Wave-U-Net with 7.25 times more parameters.

📄 PDF Abstract BibTeX arXiv:1908.06468

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Deep Residual-Dense Lattice Network for Speech Enhancement

2020-02-27 · Mohammad Nikzad, Aaron Nicolson, Yongsheng Gao, Jun Zhou 외

Convolutional neural networks (CNNs) with residual links (ResNets) and causal dilated convolutional units have been the network of choice for deep learning approaches to speech enhancement. While residual links improve g…

Speech Enhancement

Monaural Speech Enhancement Using a Multi-Branch Temporal Convolutional Network

2019-12-27 · Qiquan Zhang, Aaron Nicolson, Mingjiang Wang, Kuldip K. Paliwal 외

Deep learning has achieved substantial improvement on single-channel speech enhancement tasks. However, the performance of multi-layer perceptions (MLPs)-based methods is limited by the ability to capture the long-term e…

Speech Enhancement

SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation

2025-01-20 · Ziling Huang, Haixin Guan, Haoran Wei, Yanhua Long

Personalized speech enhancement (PSE) methods typically rely on pre-trained speaker verification models or self-designed speaker encoders to extract target speaker clues, guiding the PSE model in isolating the desired sp…

Speaker VerificationSpeech Enhancement

Dense CNN with Self-Attention for Time-Domain Speech Enhancement

2020-09-03 · Ashutosh Pandey, DeLiang Wang

Speech enhancement in the time domain is becoming increasingly popular in recent years, due to its capability to jointly enhance both the magnitude and the phase of speech. In this work, we propose a dense convolutional …

DecoderSpeech Enhancement

Deep Speech Enhancement for Reverberated and Noisy Signals using Wide Residual Networks

2019-01-03 · Dayana Ribas, Jorge Llombart, Antonio Miguel, Luis Vicente

This paper proposes a deep speech enhancement method which exploits the high potential of residual connections in a wide neural network architecture, a topology known as Wide Residual Network. This is supported on single…

Speech Enhancementspeech-recognitionSpeech Recognition