paper-with-me

Papers

A Flow-Based Neural Network for Time Domain Speech Enhancement

2021-06-16 · Martin Strauss, Bernd Edler

Speech enhancement involves the distinction of a target speech signal from an intrusive background. Although generative approaches using Variational Autoencoders or Generative Adversarial Networks (GANs) have increasingly been used in recent years, normalizing flow (NF) based systems are still scarse, despite their success in related fields. Thus, in this paper we propose a NF framework to directly model the enhancement process by density estimation of clean speech utterances conditioned on their noisy counterpart. The WaveGlow model from speech synthesis is adapted to enable direct enhancement of noisy utterances in time domain. In addition, we demonstrate that nonlinear input companding benefits the model performance by equalizing the distribution of input samples. Experimental evaluation on a publicly available dataset shows comparable results to current state-of-the-art GAN-based approaches, while surpassing the chosen baselines using objective evaluation metrics.

📄 PDF Abstract BibTeX arXiv:2106.09008

Code (0)

등록된 구현이 없습니다.

Tasks

Density EstimationSpeech EnhancementSpeech Synthesis

Methods 이 논문이 사용한 방법론

Invertible 1x1 Convolution The Invertible 1x1 Convolution is a type of convolution used in flow-based generative models that reverses the ordering of…
Affine Coupling 설명 없음
Normalizing Flows Normalizing Flows are a method for constructing complex distributions by transforming a probability density through a series of invertible mappings. By repeatedly applying…
WaveGlow WaveGlow is a flow-based generative model that generates audio by sampling from a distribution. Specifically samples are taken from a zero mean spherical Gaussian with the…

Similar Papers 제목 키워드 기반

Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement

2025-08-28 · Mattias Cross, Anton Ragni arxiv

Current flow-based generative speech enhancement methods learn curved probability paths which model a mapping between clean and noisy speech. Despite impressive performance, the implications of curved probability paths a…

Speech Enhancement

TAPLoss: A Temporal Acoustic Parameter Loss for Speech Enhancement

2023-02-16 · Yunyang Zeng, Joseph Konan, Shuo Han, David Bick 외

Speech enhancement models have greatly progressed in recent years, but still show limits in perceptual quality of their speech outputs. We propose an objective for perceptual quality based on temporal acoustic parameters…

Speaker RecognitionSpeech Enhancement

PAAPLoss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement

2023-02-16 · Muqiao Yang, Joseph Konan, David Bick, Yunyang Zeng 외

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in …

Speech EnhancementTime SeriesTime Series Analysis

A non-causal FFTNet architecture for speech enhancement

2020-06-08 · Muhammed PV Shifas, Nagaraj Adiga, Vassilis Tsiaras, Yannis Stylianou

In this paper, we suggest a new parallel, non-causal and shallow waveform domain architecture for speech enhancement based on FFTNet, a neural network for generating high quality audio waveform. In contrast to other wave…

Speech Enhancement

Neural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated Full- and Sub-Band Modeling

2023-04-18 · Zhong-Qiu Wang, Samuele Cornell, Shukjae Choi, Younglo Lee 외

We propose FSB-LSTM, a novel long short-term memory (LSTM) based architecture that integrates full- and sub-band (FSB) modeling, for single- and multi-channel speech enhancement in the short-time Fourier transform (STFT)…

Speech Enhancement