paper-with-me

Papers

Complex spectrogram enhancement by convolutional neural network with multi-metrics learning

2017-04-27 · Szu-Wei Fu, Ting-yao Hu, Yu Tsao, Xugang Lu

This paper aims to address two issues existing in the current speech enhancement methods: 1) the difficulty of phase estimations; 2) a single objective function cannot consider multiple metrics simultaneously. To solve the first problem, we propose a novel convolutional neural network (CNN) model for complex spectrogram enhancement, namely estimating clean real and imaginary (RI) spectrograms from noisy ones. The reconstructed RI spectrograms are directly used to synthesize enhanced speech waveforms. In addition, since log-power spectrogram (LPS) can be represented as a function of RI spectrograms, its reconstruction is also considered as another target. Thus a unified objective function, which combines these two targets (reconstruction of RI spectrograms and LPS), is equivalent to simultaneously optimizing two commonly used objective metrics: segmental signal-to-noise ratio (SSNR) and logspectral distortion (LSD). Therefore, the learning process is called multi-metrics learning (MML). Experimental results confirm the effectiveness of the proposed CNN with RI spectrograms and MML in terms of improved standardized evaluation metrics on a speech enhancement task.

📄 PDF Abstract BibTeX arXiv:1704.08504

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features

2024-09-19 · Younghoo Kwon, Jung-Woo Choi

We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based d…

Speech Enhancement

Phase-aware Speech Enhancement with Deep Complex U-Net

2019-03-07 · ICLR 2019 5 · Hyeong-Seok Choi, Jang-Hyun Kim, Jaesung Huh, Adrian Kim 외

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimat…

Speech Enhancementvalid

End-to-End Model for Speech Enhancement by Consistent Spectrogram Masking

2019-01-02 · Xingjian Du, Mengyao Zhu, Xuan Shi, Xinpeng Zhang 외

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier tra…

Speech Enhancement

Speech enhancement based on the integration of fully convolutional network, temporal lowpass filtering and spectrogram masking

2019-10-01 · ROCLING 2019 10 · Kuan-Yi Liu, Syu-Siang Wang, Yu Tsao, Jeih-weih Hung
Speech Enhancement

RHR-Net: A Residual Hourglass Recurrent Neural Network for Speech Enhancement

2019-04-15 · Jalal Abdulbaqi, Yue Gu, Ivan Marsic

Most current speech enhancement models use spectrogram features that require an expensive transformation and result in phase information loss. Previous work has overcome these issues by using convolutional networks to le…

Speech Enhancement