paper-with-me

홈 › Papers

Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head Attention

2022-05-03 · Xinmeng Xu, Rongzhi Gu, Yuexian Zou

Hand-crafted spatial features, such as inter-channel intensity difference (IID) and inter-channel phase difference (IPD), play a fundamental role in recent deep learning based dual-microphone speech enhancement (DMSE) systems. However, learning the mutual relationship between artificially designed spatial and spectral features is hard in the end-to-end DMSE. In this work, a novel architecture for DMSE using a multi-head cross-attention based convolutional recurrent network (MHCA-CRN) is presented. The proposed MHCA-CRN model includes a channel-wise encoding structure for preserving intra-channel features and a multi-head cross-attention mechanism for fully exploiting cross-channel features. In addition, the proposed approach specifically formulates the decoder with an extra SNR estimator to estimate frame-level SNR under a multi-task learning framework, which is expected to avoid speech distortion led by end-to-end DMSE module. Finally, a spectral gain function is adopted to further suppress the unnatural residual noise. Experiment results demonstrated superior performance of the proposed model against several state-of-the-art models.

📄 PDF Abstract BibTeX arXiv:2205.01280

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderMulti-Task LearningSpeech Enhancement

Similar Papers 제목 키워드 기반

Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario

2024-12-24 · Wen Wen, Qiang Zhou, Yu Xi, Haoyu Li 외

In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, es…

Speech Enhancement

A Causal U-net based Neural Beamforming Network for Real-Time Multi-Channel Speech Enhancement

2021-08-01 · INTERSPEECH 2021 2021 8 · Xinlei Ren, Xu Zhang, LianWu Chen, Xiguang Zheng 외

People are meeting through video conferencing more often. While single channel speech enhancement techniques are useful for the individual participants, the speech quality will be significantly degraded in large meeting …

CPUSpeech Enhancement

Microphone Array Generalization for Multichannel Narrowband Deep Speech Enhancement

2021-07-27 · Siyuan Zhang, Xiaofei Li

This paper addresses the problem of microphone array generalization for deep-learning-based end-to-end multichannel speech enhancement. We aim to train a unique deep neural network (DNN) potentially performing well on un…

Speech Enhancement

Multichannel Speech Enhancement by Raw Waveform-mapping using Fully Convolutional Networks

2019-09-26 · Chang-Le Liu, Sze-Wei Fu, You-Jin Li, Jen-Wei Huang 외

In recent years, waveform-mapping-based speech enhancement (SE) methods have garnered significant attention. These methods generally use a deep learning model to directly process and reconstruct speech waveforms. Because…

DenoisingSpeech Enhancement

Multi-channel target speech enhancement based on ERB-scaled spatial coherence features

2022-07-17 · Yicheng Hsu, Yonghan Lee, Mingsian R. Bai

Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageou…

Speech Enhancement