paper-with-me

Papers

Deep neural network Based Low-latency Speech Separation with Asymmetric analysis-Synthesis Window Pair

2021-06-22 · Shanshan Wang, Gaurav Naithani, Archontis Politis, Tuomas Virtanen

Time-frequency masking or spectrum prediction computed via short symmetric windows are commonly used in low-latency deep neural network (DNN) based source separation. In this paper, we propose the usage of an asymmetric analysis-synthesis window pair which allows for training with targets with better frequency resolution, while retaining the low-latency during inference suitable for real-time speech enhancement or assisted hearing applications. In order to assess our approach across various model types and datasets, we evaluate it with both speaker-independent deep clustering (DC) model and a speaker-dependent mask inference (MI) model. We report an improvement in separation performance of up to 1.5 dB in terms of source-to-distortion ratio (SDR) while maintaining an algorithmic latency of 8 ms.

📄 PDF Abstract BibTeX arXiv:2106.11794

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDeep ClusteringSpeech EnhancementSpeech Separation

Similar Papers 제목 키워드 기반

VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition

2020-09-09 · Quan Wang, Ignacio Lopez Moreno, Mert Saglam, Kevin Wilson 외

We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speech recognition system. Delivering such a …

CPUspeech-recognitionSpeech Recognition

Low-Latency Speech Separation Guided Diarization for Telephone Conversations

2022-04-05 · Giovanni Morrone, Samuele Cornell, Desh Raj, Luca Serafini 외

In this paper, we carry out an analysis on the use of speech separation guided diarization (SSGD) in telephone conversations. SSGD performs diarization by separating the speakers signals and then applying voice activity …

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+3

Utterance-level Permutation Invariant Training with Latency-controlled BLSTM for Single-channel Multi-talker Speech Separation

2019-12-25 · Lu Huang, Gaofeng Cheng, Pengyuan Zhang, Yi Yang 외

Utterance-level permutation invariant training (uPIT) has achieved promising progress on single-channel multi-talker speech separation task. Long short-term memory (LSTM) and bidirectional LSTM (BLSTM) are widely used as…

Speech Separation

SkiM: Skipping Memory LSTM for Low-Latency Real-Time Continuous Speech Separation

2022-01-26 · Chenda Li, Lei Yang, Weiqin Wang, Yanmin Qian

Continuous speech separation for meeting pre-processing has recently become a focused research topic. Compared to the data in utterance-level speech separation, the meeting-style audio stream lasts longer, has an uncerta…

Speech Separation

Universal Sound Separation

2019-05-08 · Ilya Kavalerov, Scott Wisdom, Hakan Erdogan, Brian Patton 외

Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of differe…

Speech EnhancementSpeech Separation