DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech Enhancement
In this study, we propose a dense frequency-time attentive network (DeFT-AN) for multichannel speech enhancement. DeFT-AN is a mask estimation network that predicts a complex spectral masking pattern for suppressing the noise and reverberation embedded in the short-time Fourier transform (STFT) of an input signal. The proposed mask estimation network incorporates three different types of blocks for aggregating information in the spatial, spectral, and temporal dimensions. It utilizes a spectral transformer with a modified feed-forward network and a temporal conformer with sequential dilated convolutions. The use of dense blocks and transformers dedicated to the three different characteristics of audio signals enables more comprehensive enhancement in noisy and reverberant environments. The remarkable performance of DeFT-AN over state-of-the-art multichannel models is demonstrated based on two popular noisy and reverberant datasets in terms of various metrics for speech quality and intelligibility.
Code (1)
Tasks
DenoisingSpeech DereverberationSpeech EnhancementSimilar Papers 제목 키워드 기반
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
This paper presents a framework for universal sound separation and polyphonic audio classification, addressing the challenges of separating and classifying individual sound sources in a multichannel mixture. The proposed…
Audio ClassificationClassificationMambaDeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
In this work, we present DeFTAN-II, an efficient multichannel speech enhancement model based on transformer architecture and subgroup processing. Despite the success of transformers in speech enhancement, they face chall…
Speech EnhancementChannel-Attention Dense U-Net for Multichannel Speech Enhancement
Supervised deep learning has gained significant attention for speech enhancement recently. The state-of-the-art deep learning methods perform the task by learning a ratio/binary mask that is applied to the mixture in the…
Speech EnhancementLSE au DEFT 2018 : Classification de tweets bas\'ee sur les r\'eseaux de neurones profonds (LSE at DEFT 2018 : Sentiment analysis model based on deep learning)
Dans ce papier, nous d{\'e}crivons les syst{\`e}mes d{\'e}velopp{\'e}s au LSE pour le DEFT 2018 sur les t{\^a}ches 1 et 2 qui consistent {\`a} classifier des tweets. La premi{\`e}re t{\^a}che consiste {\`a} d{\'e}termine…
Sentiment AnalysisConsistent ICA: Determined BSS meets spectrogram consistency
Multichannel audio blind source separation (BSS) in the determined situation (the number of microphones is equal to that of the sources), or determined BSS, is performed by multichannel linear filtering in the time-frequ…
blind source separation