A Modulation-Domain Loss for Neural-Network-based Real-time Speech Enhancement
We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learned STRFs were then used to calculate a weighted mean-squared error (MSE) in the modulation domain for training a speech enhancement system. Experiments showed that adding the modulation-domain MSE to the MSE in the spectro-temporal domain substantially improved the objective prediction of speech quality and intelligibility for real-time speech enhancement systems without incurring additional computation during inference.
Code (1)
Tasks
Speaker IdentificationSpeech DenoisingSpeech EnhancementSimilar Papers 제목 키워드 기반
Self-supervised Learning with Speech Modulation Dropout
We show that training a multi-headed self-attention-based deep network to predict deleted, information-dense 2-8 Hz speech modulations over a 1.5-second section of a speech utterance is an effective way to make machines …
Automatic Speech RecognitionSelf-Supervised Learningspeech-recognitionSpeech RecognitionOrthogonal Delay-Doppler Division Multiplexing Modulation with Tomlinson-Harashima Precoding
The orthogonal delay-Doppler (DD) division multiplexing(ODDM) modulation has been recently proposed as a promising modulation scheme for next-generation communication systems with high mobility. Despite its benefits, ODD…
The Future of Prosody: It's about Time
Prosody is usually defined in terms of the three distinct but interacting domains of pitch, intensity and duration patterning, or, more generally, as phonological and phonetic properties of 'suprasegmentals', speech segm…
Open Set Modulation Recognition Based on Dual-Channel LSTM Model
Deep neural networks have achieved great success in computer vision, speech recognition and many other areas. The potential of recurrent neural networks especially the Long Short-Term Memory (LSTM) for open set communica…
speech-recognitionSpeech RecognitionStarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion
Non-parallel multi-domain voice conversion (VC) is a technique for learning mappings among multiple domains without relying on parallel data. This is important but challenging owing to the requirement of learning multipl…
Voice Conversion