Speech Denoising with Auditory Models
Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the possibility of using deep feature representations as 'perceptual' losses with which to train denoising systems. We explored their utility by first training deep neural networks to classify either spoken words or environmental sounds from audio. We then trained an audio transform to map noisy speech to an audio waveform that minimized the difference in the deep feature representations between the output audio and the corresponding clean audio. The resulting transforms removed noise substantially better than baseline methods trained to reconstruct clean waveforms, and also outperformed previous methods using deep feature losses. However, a similar benefit was obtained simply by using losses derived from the filter bank inputs to the deep networks. The results show that deep features can guide speech enhancement, but suggest that they do not yet outperform simple alternatives that do not involve learned features.
Code (1)
Tasks
DenoisingSpeech DenoisingSpeech EnhancementSimilar Papers 제목 키워드 기반
Speech Driven Video Editing via an Audio-Conditioned Diffusion Model
Taking inspiration from recent developments in visual generative tasks using diffusion models, we propose a method for end-to-end speech-driven video editing using a denoising diffusion model. Given a video of a talking …
DenoisingFace ModelLip ReadingVideo EditingAn efficient and perceptually motivated auditory neural encoding and decoding algorithm for spiking neural networks
Auditory front-end is an integral part of a spiking neural network (SNN) when performing auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reco…
Benchmarkingspeech-recognitionSpeech RecognitionAuditory-Based Data Augmentation for End-to-End Automatic Speech Recognition
End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human au…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1Jointly Learning Visual and Auditory Speech Representations from Raw Data
We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…
Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1Contribution of Coincidence Detection to Speech Segregation in Noisy Environments
This study introduces a biologically-inspired model designed to examine the role of coincidence detection cells in speech segregation tasks. The model consists of three stages: a time-domain cochlear model that generates…