paper-with-me

Papers

Deep learning for monaural speech separation

2014-05-04 · ICASSP 2014 5 · Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, Paris Smaragdis

Monaural source separation is useful for many real-world applications though it is a challenging problem. In this paper, we study deep learning for monaural speech separation. We propose the joint optimization of the deep learning models (deep neural networks and recurrent neural networks) with an extra masking layer, which enforces a reconstruction constraint. Moreover, we explore a discriminative training criterion for the neural networks to further enhance the separation performance. We evaluate our approaches using the TIMIT speech corpus for a monaural speech separation task. Our proposed models achieve about 3.8~4.9 dB SIR gain compared to NMF models, while maintaining better SDRs and SARs.

📄 PDF Abstract BibTeX

Code (1)

posenhuang/deeplearningsourceseparation

Tasks

Deep LearningMulti-Speaker Source SeparationSpeech Separation

Similar Papers 제목 키워드 기반

Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation

2015-02-13 · Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, Paris Smaragdis

Monaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possi…

DenoisingSpeech DenoisingSpeech Separation

Monaural Multi-Speaker Speech Separation Using Efficient Transformer Model

2023-07-29 · S. Rijal, R. Neupane, S. P. Mainali, S. K. Regmi 외

Cocktail party problem is the scenario where it is difficult to separate or distinguish individual speaker from a mixed speech from several speakers. There have been several researches going on in this field but the size…

Computational EfficiencySpeech Separation

End-to-End Monaural Multi-speaker ASR System without Pretraining

2018-11-05 · Xuankai Chang, Yanmin Qian, Kai Yu, Shinji Watanabe

Recently, end-to-end models have become a popular approach as an alternative to traditional hybrid models in automatic speech recognition (ASR). The multi-speaker speech separation and recognition task is a central task …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Deep neural network techniques for monaural speech enhancement: state of the art analysis

2022-12-01 · Peter Ochieng

Deep neural networks (DNN) techniques have become pervasive in domains such as natural language processing and computer vision. They have achieved great success in these domains in task such as machine translation and im…

Art AnalysisImage GenerationMachine TranslationSpeaker Separation+2

Monaural source separation: From anechoic to reverberant environments

2021-11-15 · Tobias Cord-Landwehr, Christoph Boeddeker, Thilo von Neumann, Catalin Zorila 외

Impressive progress in neural network-based single-channel speech source separation has been made in recent years. But those improvements have been mostly reported on anechoic data, a situation that is hardly met in prac…