paper-with-me

Papers

Multi-Loss Convolutional Network with Time-Frequency Attention for Speech Enhancement

2023-06-15 · Liang Wan, Hongqing Liu, Yi Zhou, Jie Ji

The Dual-Path Convolution Recurrent Network (DPCRN) was proposed to effectively exploit time-frequency domain information. By combining the DPRNN module with Convolution Recurrent Network (CRN), the DPCRN obtained a promising performance in speech separation with a limited model size. In this paper, we explore self-attention in the DPCRN module and design a model called Multi-Loss Convolutional Network with Time-Frequency Attention(MNTFA) for speech enhancement. We use self-attention modules to exploit the long-time information, where the intra-chunk self-attentions are used to model the spectrum pattern and the inter-chunk self-attention are used to model the dependence between consecutive frames. Compared to DPRNN, axial self-attention greatly reduces the need for memory and computation, which is more suitable for long sequences of speech signals. In addition, we propose a joint training method of a multi-resolution STFT loss and a WavLM loss using a pre-trained WavLM network. Experiments show that with only 0.23M parameters, the proposed model achieves a better performance than DPCRN.

📄 PDF Abstract BibTeX arXiv:2306.08956

Code (0)

등록된 구현이 없습니다.

Tasks

Speech EnhancementSpeech Separation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses

2021-02-03 · Shengkui Zhao, Trung Hieu Nguyen, Bin Ma

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures wi…

DecoderSpeech DenoisingSpeech Enhancement

Convolutional Recurrent Neural Network with Attention for 3D Speech Enhancement

2023-06-08 · Han Yin, Jisheng Bai, Mou Wang, Siwei Huang 외

3D speech enhancement can effectively improve the auditory experience and plays a crucial role in augmented reality technology. However, traditional convolutional-based speech enhancement methods have limitations in extr…

DenoisingSpeech Enhancement

Spec2VolCAMU-Net: A Spectrogram-to-Volume Model for EEG-to-fMRI Reconstruction based on Multi-directional Time-Frequency Convolutional Attention Encoder and Vision-Mamba U-Net

2025-05-14 · Dongyi He, Shiyang Li, Bin Jiang, He Yan

High-resolution functional magnetic resonance imaging (fMRI) is essential for mapping human brain activity; however, it remains costly and logistically challenging. If comparable volumes could be generated directly from …

EEGMambaSSIM

Multi-dimensional frequency dynamic convolution with confident mean teacher for sound event detection

2023-02-18 · Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang

Recently, convolutional neural networks (CNNs) have been widely used in sound event detection (SED). However, traditional convolution is deficient in learning time-frequency domain representation of different sound event…

Event DetectionSound Event Detection

Skeleton-Based Action Recognition with Synchronous Local and Non-local Spatio-temporal Learning and Frequency Attention

2018-11-10 · Guyue Hu, Bo Cui, Shan Yu

Benefiting from its succinctness and robustness, skeleton-based action recognition has recently attracted much attention. Most existing methods utilize local networks (e.g., recurrent, convolutional, and graph convolutio…

Action RecognitionSkeleton Based Action RecognitionTemporal Action Localization