paper-with-me

홈 › Papers

Receptive Field Analysis of Temporal Convolutional Networks for Monaural Speech Dereverberation

2022-04-13 · William Ravenscroft, Stefan Goetze, Thomas Hain

Speech dereverberation is often an important requirement in robust speech processing tasks. Supervised deep learning (DL) models give state-of-the-art performance for single-channel speech dereverberation. Temporal convolutional networks (TCNs) are commonly used for sequence modelling in speech enhancement tasks. A feature of TCNs is that they have a receptive field (RF) dependent on the specific model configuration which determines the number of input frames that can be observed to produce an individual output frame. It has been shown that TCNs are capable of performing dereverberation of simulated speech data, however a thorough analysis, especially with focus on the RF is yet lacking in the literature. This paper analyses dereverberation performance depending on the model size and the RF of TCNs. Experiments using the WHAMR corpus which is extended to include room impulse responses (RIRs) with larger T60 values demonstrate that a larger RF can have significant improvement in performance when training smaller TCN models. It is also demonstrated that TCNs benefit from a wider RF when dereverberating RIRs with larger RT60 values.

📄 PDF Abstract BibTeX arXiv:2204.06439

Code (1)

jwr1995/whamr_ext 공식 구현

Tasks

Speech DereverberationSpeech Enhancement

Similar Papers 제목 키워드 기반

Monaural Speech Enhancement Using a Multi-Branch Temporal Convolutional Network

2019-12-27 · Qiquan Zhang, Aaron Nicolson, Mingjiang Wang, Kuldip K. Paliwal 외

Deep learning has achieved substantial improvement on single-channel speech enhancement tasks. However, the performance of multi-layer perceptions (MLPs)-based methods is limited by the ability to capture the long-term e…

Speech Enhancement

Utterance Weighted Multi-Dilation Temporal Convolutional Networks for Monaural Speech Dereverberation

2022-05-17 · William Ravenscroft, Stefan Goetze, Thomas Hain

Speech dereverberation is an important stage in many speech technology applications. Recent work in this area has been dominated by deep neural network models. Temporal convolutional networks (TCNs) are deep learning mod…

Speech Dereverberation

Deformable Temporal Convolutional Networks for Monaural Noisy Reverberant Speech Separation

2022-10-27 · William Ravenscroft, Stefan Goetze, Thomas Hain

Speech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation…

Speech DereverberationSpeech Separation

Video BagNet: short temporal receptive fields increase robustness in long-term action recognition

2023-08-22 · Ombretta Strafforello, Xin Liu, Klamer Schutte, Jan van Gemert

Previous work on long-term video action recognition relies on deep 3D-convolutional models that have a large temporal receptive field (RF). We argue that these models are not always the best choice for temporal modeling …

Action RecognitionTemporal Action Localization

Multi-Resolution Fully Convolutional Neural Networks for Monaural Audio Source Separation

2017-10-28 · Emad M. Grais, Hagen Wierstorf, Dominic Ward, Mark D. Plumbley

In deep neural networks with convolutional layers, each layer typically has fixed-size/single-resolution receptive field (RF). Convolutional layers with a large RF capture global information from the input features, whil…

Audio Source Separation