Deep Learning for Audio Signal Processing
Given the recent surge in developments of deep learning, this article provides a review of the state-of-the-art deep learning techniques for audio signal processing. Speech, music, and environmental sound processing are considered side-by-side, in order to point out similarities and differences between the domains, highlighting general methods, problems, key references, and potential for cross-fertilization between areas. The dominant feature representations (in particular, log-mel spectra and raw waveform) and deep learning models are reviewed, including convolutional neural networks, variants of the long short-term memory architecture, as well as more audio-specific neural network models. Subsequently, prominent deep learning application areas are covered, i.e. audio recognition (automatic speech recognition, music information retrieval, environmental sound detection, localization and tracking) and synthesis and transformation (source separation, audio enhancement, generative models for speech, sound, and music synthesis). Finally, key issues and future questions regarding deep learning applied to audio signal processing are identified.
Code (1)
Tasks
Audio Signal ProcessingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep LearningInformation RetrievalMusic Information RetrievalRetrievalspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Utterance Clustering Using Stereo Audio Channels
Utterance clustering is one of the actively researched topics in audio signal processing and machine learning. This study aims to improve the performance of utterance clustering by processing multichannel (stereo) audio …
Audio Signal ProcessingClusteringSpeaker DiarizationBridging Biological Hearing and Neuromorphic Computing: End-to-End Time-Domain Audio Signal Processing with Reservoir Computing
Despite the advancements in cutting-edge technologies, audio signal processing continues to pose challenges and lacks the precision of a human speech processing system. To address these challenges, we propose a novel app…
Speech RecognitionMulticlass Language Identification using Deep Learning on Spectral Images of Audio Signals
The first step in any voice recognition software is to determine what language a speaker is using, and ideally this process would be automated. The technique described in this paper, language identification for audio spe…
ClassificationGeneral ClassificationLanguage IdentificationMulti-class ClassificationSimulating the DFT Algorithm for Audio Processing
Since the evolution of digital computers, the storage of data has always been in terms of discrete bits that can store values of either 1 or 0. Hence, all computer programs (such as MATLAB), convert any input continuous …
A Comparison of Audio Signal Preprocessing Methods for Deep Neural Networks on Music Tagging
In this paper, we empirically investigate the effect of audio preprocessing on music tagging with deep neural networks. We perform comprehensive experiments involving audio preprocessing using different time-frequency re…
Music Tagging