paper-with-me

홈 › Papers

Deep speech inpainting of time-frequency masks

2019-10-20 · Mikolaj Kegler, Pierre Beckmann, Milos Cernak

Transient loud intrusions, often occurring in noisy environments, can completely overpower speech signal and lead to an inevitable loss of information. While existing algorithms for noise suppression can yield impressive results, their efficacy remains limited for very low signal-to-noise ratios or when parts of the signal are missing. To address these limitations, here we propose an end-to-end framework for speech inpainting, the context-based retrieval of missing or severely distorted parts of time-frequency representation of speech. The framework is based on a convolutional U-Net trained via deep feature losses, obtained using speechVGG, a deep speech feature extractor pre-trained on an auxiliary word classification task. Our evaluation results demonstrate that the proposed framework can recover large portions of missing or distorted time-frequency representation of speech, up to 400 ms and 3.2 kHz in bandwidth. In particular, our approach provided a substantial increase in STOI & PESQ objective metrics of the initially corrupted speech samples. Notably, using deep feature losses to train the framework led to the best results, as compared to conventional approaches.

📄 PDF Abstract BibTeX arXiv:1910.09058

Code (2)

MKegler/SpeechVGG 공식 구현 tf
bepierre/SpeechVGG 공식 구현 tf

Tasks

Retrieval

Methods 이 논문이 사용한 방법론

Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

IE-NeRF: Inpainting Enhanced Neural Radiance Fields in the Wild

2024-07-15 · Shuaixian Wang, Haoran Xu, Yaokun Li, Jiwei Chen 외

We present a novel approach for synthesizing realistic novel views using Neural Radiance Fields (NeRF) with uncontrolled photos in the wild. While NeRF has shown impressive results in controlled settings, it struggles wi…

Image InpaintingNeRF

GLaMa: Joint Spatial and Frequency Loss for General Image Inpainting

2022-05-15 · Zeyu Lu, Junjun Jiang, Junqin Huang, Gang Wu 외

The purpose of image inpainting is to recover scratches and damaged areas using context information from remaining parts. In recent years, thanks to the resurgence of convolutional neural networks (CNNs), image inpaintin…

Image InpaintingSSIM

Iterative Image Inpainting with Structural Similarity Mask for Anomaly Detection

2021-01-01 · Hitoshi Nakanishi, Masahiro Suzuki, Yutaka Matsuo

Autoencoders have emerged as popular methods for unsupervised anomaly detection. Autoencoders trained on the normal data are expected to reconstruct only the normal features, allowing anomaly detection by thresholding re…

Anomaly DetectionImage InpaintingUnsupervised Anomaly Detection

Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation

2018-09-20 · Yi Luo, Nima Mesgarani

Single-channel, speaker-independent speech separation methods have recently seen great progress. However, the accuracy, latency, and computational cost of such methods remain insufficient. The majority of the previous me…

Multi-task Audio Source SeperationMusic Source SeparationSpeaker SeparationSpeech Enhancement+1

Enhancement of Spatial Clustering-Based Time-Frequency Masks using LSTM Neural Networks

2020-12-02 · Felix Grezes, Zhaoheng Ni, Viet Anh Trinh, Michael Mandel

Recent works have shown that Deep Recurrent Neural Networks using the LSTM architecture can achieve strong single-channel speech enhancement by estimating time-frequency masks. However, these models do not naturally gene…

ClusteringSpeech Enhancement