Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features of self-supervised learning (SSL) models. This work is the first to incorporate SSL features with causality into an SE model. The causal SSL features are encoded and combined with spectrogram features using feature-wise linear modulation to estimate a mask for enhancing the noisy input speech. Simultaneously, we quantize the causal SSL features using vector quantization to represent phonetic characteristics as semantic tokens. The model not only encodes SSL features but also predicts the future semantic tokens in multi-task learning (MTL). The experimental results using VoiceBank + DEMAND dataset show that our proposed method achieves 2.88 in PESQ, especially with semantic prediction MTL, in which we confirm that the semantic prediction played an important role in causal SE.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Task LearningQuantizationSelf-Supervised LearningSpeech EnhancementSimilar Papers 제목 키워드 기반
A study on speech enhancement using exponent-only floating point quantized neural network (EOFP-QNN)
Numerous studies have investigated the effectiveness of neural network quantization on pattern classification tasks. The present study, for the first time, investigated the performance of speech enhancement (a regression…
QuantizationregressionSpeech EnhancementInference and Denoise: Causal Inference-based Neural Speech Enhancement
This study addresses the speech enhancement (SE) task within the causal inference paradigm by modeling the noise presence as an intervention. Based on the potential outcome framework, the proposed causal inference-based …
Causal InferenceSpeech EnhancementIncorporating Real-world Noisy Speech in Neural-network-based Speech Enhancement Systems
Supervised speech enhancement relies on parallel databases of degraded speech signals and their clean reference signals during training. This setting prohibits the use of real-world degraded speech data that may better r…
Speech EnhancementTripletA non-causal FFTNet architecture for speech enhancement
In this paper, we suggest a new parallel, non-causal and shallow waveform domain architecture for speech enhancement based on FFTNet, a neural network for generating high quality audio waveform. In contrast to other wave…
Speech EnhancementCausal Signal-Based DCCRN with Overlapped-Frame Prediction for Online Speech Enhancement
The aim of speech enhancement is to improve speech signal quality and intelligibility from a noisy microphone signal. In many applications, it is crucial to enable processing with small computational complexity and minim…
Speech Enhancement