paper-with-me

Papers

Multi-View Attention Transfer for Efficient Speech Enhancement

2022-08-22 · WooSeok Shin, Hyun Joon Park, Jin Sob Kim, Byung Hoon Lee, Sung Won Han

Recent deep learning models have achieved high performance in speech enhancement; however, it is still challenging to obtain a fast and low-complexity model without significant performance degradation. Previous knowledge distillation studies on speech enhancement could not solve this problem because their output distillation methods do not fit the speech enhancement task in some aspects. In this study, we propose multi-view attention transfer (MV-AT), a feature-based distillation, to obtain efficient speech enhancement models in the time domain. Based on the multi-view features extraction model, MV-AT transfers multi-view knowledge of the teacher network to the student network without additional parameters. The experimental results show that the proposed method consistently improved the performance of student models of various sizes on the Valentini and deep noise suppression (DNS) datasets. MANNER-S-8.1GF with our proposed method, a lightweight model for efficient deployment, achieved 15.4x and 4.71x fewer parameters and floating-point operations (FLOPs), respectively, compared to the baseline model with similar performance.

📄 PDF Abstract BibTeX arXiv:2208.10367

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationSpeech Enhancement

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Multi-view Attention-based Speech Enhancement Model for Noise-robust Automatic Speech Recognition

2020-09-01 · ROCLING 2020 9 · Fu-An Chao, Jeih-weih Hung, Berlin Chen
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

MANNER: Multi-view Attention Network for Noise Erasure

2022-03-04 · Hyun Joon Park, Byung Ha Kang, WooSeok Shin, Jin Sob Kim 외

In the field of speech enhancement, time domain methods have difficulties in achieving both high performance and efficiency. Recently, dual-path models have been adopted to represent long sequential features, but they st…

DecoderSpeech Enhancement

Deep neural network techniques for monaural speech enhancement: state of the art analysis

2022-12-01 · Peter Ochieng

Deep neural networks (DNN) techniques have become pervasive in domains such as natural language processing and computer vision. They have achieved great success in these domains in task such as machine translation and im…

Art AnalysisImage GenerationMachine TranslationSpeaker Separation+2

U-Former: Improving Monaural Speech Enhancement with Multi-head Self and Cross Attention

2022-05-18 · Xinmeng Xu, Jianjun Hao

For supervised speech enhancement, contextual information is important for accurate spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term c…

DecoderSpeech Enhancement

BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding

2024-08-13 · Alimjan Mattursun, Liejun Wang, Yinfeng Yu

Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportu…

DenoisingSelf-Supervised LearningSpeech Enhancement