paper-with-me

Papers

Supervised Speech Separation Based on Deep Learning: An Overview

2017-08-24 · DeLiang Wang, Jitong Chen

Speech separation is the task of separating target speech from background interference. Traditionally, speech separation is studied as a signal processing problem. A more recent approach formulates speech separation as a supervised learning problem, where the discriminative patterns of speech, speakers, and background noise are learned from training data. Over the past decade, many supervised separation algorithms have been put forward. In particular, the recent introduction of deep learning to supervised speech separation has dramatically accelerated progress and boosted separation performance. This article provides a comprehensive overview of the research on deep learning based supervised speech separation in the last several years. We first introduce the background of speech separation and the formulation of supervised separation. Then we discuss three main components of supervised separation: learning machines, training targets, and acoustic features. Much of the overview is on separation algorithms where we review monaural methods, including speech enhancement (speech-nonspeech separation), speaker separation (multi-talker separation), and speech dereverberation, as well as multi-microphone techniques. The important issue of generalization, unique to supervised learning, is discussed. This overview provides a historical perspective on how advances are made. In addition, we discuss a number of conceptual issues, including what constitutes the target source.

📄 PDF Abstract BibTeX arXiv:1708.07524

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningSpeaker SeparationSpeech DereverberationSpeech EnhancementSpeech Separation

Similar Papers 제목 키워드 기반

An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation

2020-08-21 · Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu 외

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, th…

Deep LearningSpeech EnhancementSpeech Separation

End-to-end training of time domain audio separation and recognition

2019-12-18 · Thilo von Neumann, Keisuke Kinoshita, Lukas Drude, Christoph Boeddeker 외

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based t…

Speaker Recognitionspeech-recognitionSpeech RecognitionSpeech Separation

Investigating self-supervised learning for speech enhancement and separation

2022-03-15 · Zili Huang, Shinji Watanabe, Shu-wen Yang, Paola Garcia 외

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a…

Self-Supervised LearningSpeech EnhancementSpeech Separation

Using Optimal Ratio Mask as Training Target for Supervised Speech Separation

2017-09-04 · Shasha Xia, Hao Li, Xueliang Zhang

Supervised speech separation uses supervised learning algorithms to learn a mapping from an input noisy signal to an output target. With the fast development of deep learning, supervised separation has become the most im…

Speech Separation

Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training

2020-10-29 · Sung-Feng Huang, Shun-Po Chuang, Da-Rong Liu, Yi-Chen Chen 외

Speech separation has been well developed, with the very successful permutation invariant training (PIT) approach, although the frequent label assignment switching happening during PIT training remains to be a problem wh…

Speaker SeparationSpeech EnhancementSpeech Separation