paper-with-me

Papers

Speech Denoising with Auditory Models

2020-11-21 · Mark R. Saddler, Andrew Francl, Jenelle Feather, Kaizhi Qian, Yang Zhang, Josh H. McDermott

Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the possibility of using deep feature representations as 'perceptual' losses with which to train denoising systems. We explored their utility by first training deep neural networks to classify either spoken words or environmental sounds from audio. We then trained an audio transform to map noisy speech to an audio waveform that minimized the difference in the deep feature representations between the output audio and the corresponding clean audio. The resulting transforms removed noise substantially better than baseline methods trained to reconstruct clean waveforms, and also outperformed previous methods using deep feature losses. However, a similar benefit was obtained simply by using losses derived from the filter bank inputs to the deep networks. The results show that deep features can guide speech enhancement, but suggest that they do not yet outperform simple alternatives that do not involve learned features.

📄 PDF Abstract BibTeX arXiv:2011.10706

Code (1)

msaddler/auditory-model-denoising tf

Tasks

DenoisingSpeech DenoisingSpeech Enhancement

Similar Papers 제목 키워드 기반

Speech Driven Video Editing via an Audio-Conditioned Diffusion Model

2023-01-10 · Dan Bigioi, Shubhajit Basak, Michał Stypułkowski, Maciej Zięba 외

Taking inspiration from recent developments in visual generative tasks using diffusion models, we propose a method for end-to-end speech-driven video editing using a denoising diffusion model. Given a video of a talking …

DenoisingFace ModelLip ReadingVideo Editing

An efficient and perceptually motivated auditory neural encoding and decoding algorithm for spiking neural networks

2019-09-03 · Zihan Pan, Yansong Chua, Jibin Wu, Malu Zhang 외

Auditory front-end is an integral part of a spiking neural network (SNN) when performing auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reco…

Benchmarkingspeech-recognitionSpeech Recognition

Auditory-Based Data Augmentation for End-to-End Automatic Speech Recognition

2022-04-08 · Zehai Tu, Jack Deadman, Ning Ma, Jon Barker

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human au…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1

Jointly Learning Visual and Auditory Speech Representations from Raw Data

2022-12-12 · Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis 외

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1

Contribution of Coincidence Detection to Speech Segregation in Noisy Environments

2024-05-09 · Asaf Zorea, Miriam Furst

This study introduces a biologically-inspired model designed to examine the role of coincidence detection cells in speech segregation tasks. The model consists of three stages: a time-domain cochlear model that generates…