paper-with-me

홈 › Papers

Speech Enhancement with Perceptually-motivated Optimization and Dual Transformations

2022-09-24 · Xucheng Wan, Kai Liu, Ziqing Du, Huan Zhou

To address the monaural speech enhancement problem, numerous research studies have been conducted to enhance speech via operations either in time-domain on the inner-domain learned from the speech mixture or in time--frequency domain on the fixed full-band short time Fourier transform (STFT) spectrograms. Very recently, a few studies on sub-band based speech enhancement have been proposed. By enhancing speech via operations on sub-band spectrograms, those studies demonstrated competitive performances on the benchmark dataset of DNS2020. Despite attractive, this new research direction has not been fully explored and there is still room for improvement. As such, in this study, we delve into the latest research direction and propose a sub-band based speech enhancement system with perceptually-motivated optimization and dual transformations, called PT-FSE. Specially, our proposed PT-FSE model improves its backbone, a full-band and sub-band fusion model, by three efforts. First, we design a frequency transformation module that aims to strengthen the global frequency correlation. Then a temporal transformation is introduced to capture long range temporal contexts. Lastly, a novel loss, with leverage of properties of human auditory perception, is proposed to facilitate the model to focus on low frequency enhancement. To validate the effectiveness of our proposed model, extensive experiments are conducted on the DNS2020 dataset. Experimental results show that our PT-FSE system achieves substantial improvements over its backbone, but also outperforms the current state-of-the-art while being 27\% smaller than the SOTA. With average NB-PESQ of 3.57 on the benchmark dataset, our system offers the best speech enhancement results reported till date.

📄 PDF Abstract BibTeX arXiv:2209.11905

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

A two-step backward compatible fullband speech enhancement system

2022-01-26 · Xu Zhang, LianWu Chen, Xiguang Zheng, Xinlei Ren 외

Speech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system…

Speech EnhancementVocal Bursts Valence Prediction

A Perceptually-Motivated Approach for Low-Complexity, Real-Time Enhancement of Fullband Speech

2020-08-27 · Interspeech 2020 8

Over the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly i…

CPUSpeech Enhancement

DeepFilterNet: Perceptually Motivated Real-Time Speech Enhancement

2023-05-14 · Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante-B., Andreas Maier

Multi-frame algorithms for single-channel speech enhancement are able to take advantage from short-time correlations within the speech signal. Deep Filtering (DF) was proposed to directly estimate a complex filter in fre…

CPUSpeech Enhancement

Stable Training of DNN for Speech Enhancement based on Perceptually-Motivated Black-Box Cost Function

2020-02-14 · Masaki Kawanaka, Yuma Koizumi, Ryoichi Miyazaki, Kohei Yatabe

Improving subjective sound quality of enhanced signals is one of the most important missions in speech enhancement. For evaluating the subjective quality, several methods related to perceptually-motivated objective sound…

Reinforcement LearningSpeech Enhancement

Aligning Generative Speech Enhancement with Perceptual Feedback

2025-07-14 · Haoyang Li, Nana Hou, Yuchen Hu, Jixun Yao 외 arxiv

Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives that weakly reflect human perception. This …

Speech Enhancement