paper-with-me

Papers

Demystifying TasNet: A Dissecting Approach

2019-11-20 · Jens Heitkaemper, Darius Jakobeit, Christoph Boeddeker, Lukas Drude, Reinhold Haeb-Umbach

In recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments. In this paper we dissect the gains of the time-domain audio separation network (TasNet) approach by gradually replacing components of an utterance-level permutation invariant training (u-PIT) based separation system in the frequency domain until the TasNet system is reached, thus blending components of frequency domain approaches with those of time domain approaches. Some of the intermediate variants achieve comparable signal-to-distortion ratio (SDR) gains to TasNet, but retain the advantage of frequency domain processing: compatibility with classic signal processing tools such as frequency-domain beamforming and the human interpretability of the masks. Furthermore, we show that the scale invariant signal-to-distortion ratio (si-SDR) criterion used as loss function in TasNet is related to a logarithmic mean square error criterion and that it is this criterion which contributes most reliable to the performance advantage of TasNet. Finally, we critically assess which gains in a noise-free single channel environment generalize to more realistic reverberant conditions.

📄 PDF Abstract BibTeX arXiv:1911.08895

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

Beam-Guided TasNet: An Iterative Speech Separation Framework with Multi-Channel Output

2021-02-05 · Hangting Chen, Yang Yi, Dang Feng, Pengyuan Zhang

Time-domain audio separation network (TasNet) has achieved remarkable performance in blind source separation (BSS). Classic multi-channel speech processing framework employs signal estimation and beamforming. For example…

blind source separationSpeech Separation

MITAS: A Compressed Time-Domain Audio Separation Network with Parameter Sharing

2019-12-09 · Chao-I Tuan, Yuan-Kuei Wu, Hung-Yi Lee, Yu Tsao

Deep learning methods have brought substantial advancements in speech separation (SS). Nevertheless, it remains challenging to deploy deep-learning-based models on edge devices. Thus, identifying an effective way to comp…

Speech Separation

SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling

2024-07-01 · Hiroshi Sato, Takafumi Moriya, Masato Mimura, Shota Horiguchi 외

Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real-time TSE is challenging as the computat…

Target Speaker Extraction

X-TaSNet: Robust and Accurate Time-Domain Speaker Extraction Network

2020-10-24

Extracting the speech of a target speaker from mixed audios, based on a reference speech from the target speaker, is a challenging yet powerful technology in speech processing. Recent studies of speaker-independent speec…

Speech Separation

An empirical study of Conv-TasNet

2020-02-20 · Berkan Kadioglu, Michael Horgan, Xiaoyu Liu, Jordi Pons 외

Conv-TasNet is a recently proposed waveform-based deep neural network that achieves state-of-the-art performance in speech source separation. Its architecture consists of a learnable encoder/decoder and a separator that …

Decoder