paper-with-me

Papers

TRUNet: Transformer-Recurrent-U Network for Multi-channel Reverberant Sound Source Separation

2021-10-08 · Ali Aroudi, Stefan Uhlich, Marc Ferras Font

In recent years, many deep learning techniques for single-channel sound source separation have been proposed using recurrent, convolutional and transformer networks. When multiple microphones are available, spatial diversity between speakers and background noise in addition to spectro-temporal diversity can be exploited by using multi-channel filters for sound source separation. Aiming at end-to-end multi-channel source separation, in this paper we propose a transformer-recurrent-U network (TRUNet), which directly estimates multi-channel filters from multi-channel input spectra. TRUNet consists of a spatial processing network with an attention mechanism across microphone channels aiming at capturing the spatial diversity, and a spectro-temporal processing network aiming at capturing spectral and temporal diversities. In addition to multi-channel filters, we also consider estimating single-channel filters from multi-channel input spectra using TRUNet. We train the network on a large reverberant dataset using a combined compressed mean-squared error loss function, which further improves the sound separation performance. We evaluate the network on a realistic and challenging reverberant dataset, generated from measured room impulse responses of an actual microphone array. The experimental results on realistic reverberant sound source separation show that the proposed TRUNet outperforms state-of-the-art single-channel and multi-channel source separation methods.

📄 PDF Abstract BibTeX arXiv:2110.04047

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

Vision Transformers increase efficiency of 3D cardiac CT multi-label segmentation

2023-10-13 · Lee Jollans, Mariana Bustamante, Lilian Henriksson, Anders Persson 외

Accurate segmentation of the heart is essential for personalized blood flow simulations and surgical intervention planning. Segmentations need to be accurate in every spatial dimension, which is not ensured by segmenting…

3D Semantic SegmentationComputed Tomography (CT)Image SegmentationMedical Image Segmentation+2

AmbiSep: Ambisonic-to-Ambisonic Reverberant Speech Separation Using Transformer Networks

2022-06-13 · Adrian Herzog, Srikanth Raj Chetupalli, Emanuël A. P. Habets

Consider a multichannel Ambisonic recording containing a mixture of several reverberant speech signals. Retreiving the reverberant Ambisonic signals corresponding to the individual speech sources blindly from the mixture…

Speech Separation

End-to-End Multi-speaker Speech Recognition with Transformer

2020-02-10 · Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux 외

Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-channel scenarios. In this work, we explor…

Decoderspeech-recognitionSpeech Recognition

DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech Enhancement

2022-12-15 · Dongheon Lee, Jung-Woo Choi

In this study, we propose a dense frequency-time attentive network (DeFT-AN) for multichannel speech enhancement. DeFT-AN is a mask estimation network that predicts a complex spectral masking pattern for suppressing the …

DenoisingSpeech DereverberationSpeech Enhancement

The PCG-AIID System for L3DAS22 Challenge: MIMO and MISO convolutional recurrent Network for Multi Channel Speech Enhancement and Speech Recognition

2022-02-21 · Jingdong Li, Yuanyuan Zhu, Dawei Luo, Yun Liu 외

This paper described the PCG-AIID system for L3DAS22 challenge in Task 1: 3D speech enhancement in office reverberant environment. We proposed a two-stage framework to address multi-channel speech denoising and dereverbe…

DenoisingSpeech DenoisingSpeech Enhancementspeech-recognition+1