Utilizing Domain Knowledge in End-to-End Audio Processing
End-to-end neural network based approaches to audio modelling are generally outperformed by models trained on high-level data representations. In this paper we present preliminary work that shows the feasibility of training the first layers of a deep convolutional neural network (CNN) model to learn the commonly-used log-scaled mel-spectrogram transformation. Secondly, we demonstrate that upon initializing the first layers of an end-to-end CNN classifier with the learned transformation, convergence and performance on the ESC-50 environmental sound classification dataset are similar to a CNN-based model trained on the highly pre-processed log-scaled mel-spectrogram features.
Code (1)
Tasks
Environmental Sound ClassificationGeneral ClassificationSound ClassificationSimilar Papers 제목 키워드 기반
DDSP: Differentiable Digital Signal Processing
Most generative models of audio directly generate samples in one of two domains: time or frequency. While sufficient to express any signal, these representations are inefficient, as they do not utilize existing knowledge…
Audio GenerationAudio SynthesisAnnotation-free Automatic Music Transcription with Scalable Synthetic Data and Adversarial Domain Confusion
Automatic Music Transcription (AMT) is a vital technology in the field of music information processing. Despite recent enhancements in performance due to machine learning techniques, current methods typically attain high…
Music TranscriptionAudio Self-supervised Learning: A Survey
Inspired by the humans' cognitive ability to generalise knowledge and skills, Self-Supervised Learning (SSL) targets at discovering general representations from large-scale data without requiring human annotations, which…
Self-Supervised LearningSurveymllm-shap: A Shapley Value Explainability Platform for Text-Audio Multimodal Large Language Models
We introduce mllm-shap, an open-source Python framework designed to extend Shapley Value (SV) explainability from text-only Large Language Models to Multimodal LLMs (MLLMs) processing joint text and audio inputs. While t…
Automatic Depression Detection: An Emotional Audio-Textual Corpus and a GRU/BiLSTM-based Model
Depression is a global mental health problem, the worst case of which can lead to suicide. An automatic depression detection system provides great help in facilitating depression self-assessment and improving diagnostic …
Depression DetectionDiagnostic