paper-with-me

홈 › Papers

Reinforcement Learning To Adapt Speech Enhancement to Instantaneous Input Signal Quality

2017-11-29 · Rasool Fakoor, Xiaodong He, Ivan Tashev, Shuayb Zarar

Today, the optimal performance of existing noise-suppression algorithms, both data-driven and those based on classic statistical methods, is range bound to specific levels of instantaneous input signal-to-noise ratios. In this paper, we present a new approach to improve the adaptivity of such algorithms enabling them to perform robustly across a wide range of input signal and noise types. Our methodology is based on the dynamic control of algorithmic parameters via reinforcement learning. Specifically, we model the noise-suppression module as a black box, requiring no knowledge of the algorithmic mechanics except a simple feedback from the output. We utilize this feedback as the reward signal for a reinforcement-learning agent that learns a policy to adapt the algorithmic parameters for every incoming audio frame (16 ms of data). Our preliminary results show that such a control mechanism can substantially increase the overall performance of the underlying noise-suppression algorithm; 42% and 16% improvements in output SNR and MSE, respectively, when compared to no adaptivity.

📄 PDF Abstract BibTeX arXiv:1711.10791

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Speech Enhancement

Similar Papers 제목 키워드 기반

Instantaneous PSD Estimation for Speech Enhancement based on Generalized Principal Components

2020-07-01

Power spectral density (PSD) estimates of various microphone signal components are essential to many speech enhancement procedures. As speech is highly non-nonstationary, performance improvements may be gained by maintai…

Speech Enhancement

MeanFlowSE: one-step generative speech enhancement via conditional mean flow

2025-09-18 · Duojia Li, Shenghui Lu, Hongchen Pan, Zongyi Zhan 외 arxiv

Multistep inference is a bottleneck for real-time generative speech enhancement because flow- and diffusion-based systems learn an instantaneous velocity field and therefore rely on iterative ordinary differential equati…

Knowledge DistillationSpeech Enhancement

An Investigation of End-to-End Models for Robust Speech Recognition

2021-02-11 · Archiki Prasad, Preethi Jyothi, Rajbabu Velmurugan

End-to-end models for robust automatic speech recognition (ASR) have not been sufficiently well-explored in prior work. With end-to-end models, one could choose to preprocess the input speech using speech enhancement tec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMulti-Task Learning+4

A Flow-Based Neural Network for Time Domain Speech Enhancement

2021-06-16 · Martin Strauss, Bernd Edler

Speech enhancement involves the distinction of a target speech signal from an intrusive background. Although generative approaches using Variational Autoencoders or Generative Adversarial Networks (GANs) have increasingl…

Density EstimationSpeech EnhancementSpeech Synthesis

Trainable Adaptive Window Switching for Speech Enhancement

2018-11-05 · Yuma Koizumi, Noboru Harada, Yoichi Haneda

This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask proces…

Speech Enhancement