paper-with-me

Papers

MITAS: A Compressed Time-Domain Audio Separation Network with Parameter Sharing

2019-12-09 · Chao-I Tuan, Yuan-Kuei Wu, Hung-Yi Lee, Yu Tsao

Deep learning methods have brought substantial advancements in speech separation (SS). Nevertheless, it remains challenging to deploy deep-learning-based models on edge devices. Thus, identifying an effective way to compress these large models without hurting SS performance has become an important research topic. Recently, TasNet and Conv-TasNet have been proposed. They achieved state-of-the-art results on several standardized SS tasks. Moreover, their low latency natures make them definitely suitable for real-time on-device applications. In this study, we propose two parameter-sharing schemes to lower the memory consumption on TasNet and Conv-TasNet. Accordingly, we derive a novel so-called MiTAS (Mini TasNet). Our experimental results first confirmed the robustness of our MiTAS on two types of perturbations in mixed audio. We also designed a series of ablation experiments to analyze the relation between SS performance and the amount of parameters in the model. The results show that MiTAS is able to reduce the model size by a factor of four while maintaining comparable SS performance with improved stability as compared to TasNet and Conv-TasNet. This suggests that MiTAS is more suitable for real-time low latency applications.

📄 PDF Abstract BibTeX arXiv:1912.03884

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Speech Separation using Neural Audio Codecs with Embedding Loss

2024-11-27 · Jia Qi Yip, Chin Yuen Kwok, Bin Ma, Eng Siong Chng

Neural audio codecs have revolutionized audio processing by enabling speech tasks to be performed on highly compressed representations. Recent work has shown that speech separation can be achieved within these compressed…

Speech Separation

RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation

2023-09-29 · Samuel Pegg, Kai Li, Xiaolin Hu

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing stat…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Separation+1

Music Source Separation in the Waveform Domain

2019-11-27 · Alexandre Défossez, Nicolas Usunier, Léon Bottou, Francis Bach

Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any othe…

Audio GenerationAudio SynthesisData AugmentationMulti-task Audio Source Seperation+2

Sparse Gaussian Process Audio Source Separation Using Spectrum Priors in the Time-Domain

2018-10-30 · Pablo A. Alvarado, Mauricio A. Álvarez, Dan Stowell

Gaussian process (GP) audio source separation is a time-domain approach that circumvents the inherent phase approximation issue of spectrogram based methods. Furthermore, through its kernel, GPs elegantly incorporate pri…

Audio Source Separation

Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention

2021-06-17 · Efthymios Tzinis, Scott Wisdom, Tal Remez, John R. Hershey

We introduce a state-of-the-art audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify limit…

Unsupervised Pre-training