paper-with-me

홈 › Papers

Ultra-Lightweight Speech Separation via Group Communication

2020-11-18

Model size and complexity remain the biggest challenges in the deployment of speech enhancement and separation systems on low-resource devices such as earphones and hearing aids. Although methods such as compression, distillation and quantization can be applied to large models, they often come with a cost on the model performance. In this paper, we provide a simple model design paradigm that explicitly designs ultra-lightweight models without sacrificing the performance. Motivated by the sub-band frequency-LSTM (F-LSTM) architectures, we introduce the group communication (GroupComm), where a feature vector is split into smaller groups and a small processing block is used to perform inter-group communication. Unlike standard F-LSTM models where the sub-band outputs are concatenated, an ultra-small module is applied on all the groups in parallel, which allows a significant decrease on the model size. Experiment results show that comparing with a strong baseline model which is already lightweight, GroupComm can achieve on par performance with 35.6 times fewer parameters and 2.3 times fewer operations.

📄 PDF Abstract BibTeX arXiv:2011.08397

Code (0)

등록된 구현이 없습니다.

Tasks

QuantizationSpeech EnhancementSpeech Separation

Similar Papers 제목 키워드 기반

Group Communication with Context Codec for Lightweight Source Separation

2020-12-14 · Yi Luo, Cong Han, Nima Mesgarani

Despite the recent progress on neural network architectures for speech separation, the balance between the model size, model complexity and model performance is still an important and challenging problem for the deployme…

DecoderSpeech EnhancementSpeech Separation

UltraUNet: Real-Time Ultrasound Tongue Segmentation for Diverse Linguistic and Imaging Conditions

2025-09-27 · Alisher Myrgyyassov, Zhen Song, Yu Sun, Bruce Xiao Wang 외 arxiv

Ultrasound tongue imaging (UTI) is a non-invasive and cost-effective tool for studying speech articulation, motor control, and related disorders. However, real-time tongue contour segmentation remains challenging due to …

Ultra Fast Speech Separation Model with Teacher Student Learning

2022-04-27 · Sanyuan Chen, Yu Wu, Zhuo Chen, Jian Wu 외

Transformer has been successfully applied to speech separation recently with its strong long-dependency modeling capacity using a self-attention mechanism. However, Transformer tends to have heavy run-time costs due to t…

Computational EfficiencySpeech Separation

LiMuSE: Lightweight Multi-modal Speaker Extraction

2021-11-07 · Qinghua Liu, Yating Huang, Yunzhe Hao, Jiaming Xu 외

Multi-modal cues, including spatial information, facial expression and voiceprint, are introduced to the speech separation and speaker extraction tasks to serve as complementary information to achieve better performance.…

Model CompressionQuantizationSpeech Separation

Separating Long-Form Speech with Group-Wise Permutation Invariant Training

2021-10-27 · Wangyou Zhang, Zhuo Chen, Naoyuki Kanda, Shujie Liu 외

Multi-talker conversational speech processing has drawn many interests for various applications such as meeting transcription. Speech separation is often required to handle overlapped speech that is commonly observed in …

FormSpeech Separation