paper-with-me

Papers

DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech Enhancement

2022-12-15 · Dongheon Lee, Jung-Woo Choi

In this study, we propose a dense frequency-time attentive network (DeFT-AN) for multichannel speech enhancement. DeFT-AN is a mask estimation network that predicts a complex spectral masking pattern for suppressing the noise and reverberation embedded in the short-time Fourier transform (STFT) of an input signal. The proposed mask estimation network incorporates three different types of blocks for aggregating information in the spatial, spectral, and temporal dimensions. It utilizes a spectral transformer with a modified feed-forward network and a temporal conformer with sequential dilated convolutions. The use of dense blocks and transformers dedicated to the three different characteristics of audio signals enables more comprehensive enhancement in noisy and reverberant environments. The remarkable performance of DeFT-AN over state-of-the-art multichannel models is demonstrated based on two popular noisy and reverberant datasets in terms of various metrics for speech quality and intelligibility.

📄 PDF Abstract BibTeX arXiv:2212.07570

Code (1)

donghoney0416/DeFT-AN pytorch

Tasks

DenoisingSpeech DereverberationSpeech Enhancement

Similar Papers 제목 키워드 기반

DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification

2024-09-19 · Dongheon Lee, Jung-Woo Choi

This paper presents a framework for universal sound separation and polyphonic audio classification, addressing the challenges of separating and classifying individual sound sources in a multichannel mixture. The proposed…

Audio ClassificationClassificationMamba

DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing

2023-08-30 · Dongheon Lee, Jung-Woo Choi

In this work, we present DeFTAN-II, an efficient multichannel speech enhancement model based on transformer architecture and subgroup processing. Despite the success of transformers in speech enhancement, they face chall…

Speech Enhancement

Channel-Attention Dense U-Net for Multichannel Speech Enhancement

2020-01-30 · Bahareh Tolooshams, Ritwik Giri, Andrew H. Song, Umut Isik 외

Supervised deep learning has gained significant attention for speech enhancement recently. The state-of-the-art deep learning methods perform the task by learning a ratio/binary mask that is applied to the mixture in the…

Speech Enhancement

LSE au DEFT 2018 : Classification de tweets bas\'ee sur les r\'eseaux de neurones profonds (LSE at DEFT 2018 : Sentiment analysis model based on deep learning)

2018-05-01 · JEPTALNRECITAL 2018 5 · Antoine Sainson, Hugo Linsenmaier, Alex Majed, re 외

Dans ce papier, nous d{\'e}crivons les syst{\`e}mes d{\'e}velopp{\'e}s au LSE pour le DEFT 2018 sur les t{\^a}ches 1 et 2 qui consistent {\`a} classifier des tweets. La premi{\`e}re t{\^a}che consiste {\`a} d{\'e}termine…

Sentiment Analysis

Consistent ICA: Determined BSS meets spectrogram consistency

2020-05-20 · Kohei Yatabe

Multichannel audio blind source separation (BSS) in the determined situation (the number of microphones is equal to that of the sources), or determined BSS, is performed by multichannel linear filtering in the time-frequ…

blind source separation