A Training Framework for Stereo-Aware Speech Enhancement using Deep Neural Networks
Deep learning-based speech enhancement has shown unprecedented performance in recent years. The most popular mono speech enhancement frameworks are end-to-end networks mapping the noisy mixture into an estimate of the clean speech. With growing computational power and availability of multichannel microphone recordings, prior works have aimed to incorporate spatial statistics along with spectral information to boost up performance. Despite an improvement in enhancement performance of mono output, the spatial image preservation and subjective evaluations have not gained much attention in the literature. This paper proposes a novel stereo-aware framework for speech enhancement, i.e., a training loss for deep learning-based speech enhancement to preserve the spatial image while enhancing the stereo mixture. The proposed framework is model independent, hence it can be applied to any deep learning based architecture. We provide an extensive objective and subjective evaluation of the trained models through a listening test. We show that by regularizing for an image preservation loss, the overall performance is improved, and the stereo aspect of the speech is better preserved.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningSpeech EnhancementSimilar Papers 제목 키워드 기반
Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure
We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our met…
Speech EnhancementStereo Speech Enhancement Using Custom Mid-Side Signals and Monaural Processing
Speech Enhancement (SE) systems typically operate on monaural input and are used for applications including voice communications and capture cleanup for user generated content. Recent advancements and changes in the devi…
Speech EnhancementPlugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
The expectation to deploy a universal neural network for speech enhancement, with the aim of improving noise robustness across diverse speech processing tasks, faces challenges due to the existing lack of awareness withi…
Data AugmentationSpeech EnhancementSpeech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
This paper introduces a speech enhancement solution tailored for true wireless stereo (TWS) earbuds on-device usage. The solution was specifically designed to support conversations in noisy environments, with active nois…
Speech EnhancementConsistency-aware Self-Training for Iterative-based Stereo Matching
Iterative-based methods have become mainstream in stereo matching due to their high performance. However, these methods heavily rely on labeled data and face challenges with unlabeled real-world data. To this end, we pro…
Stereo Matching