Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
This study explores the field of audio classification from raw waveform using Convolutional Neural Networks (CNNs), a method that eliminates the need for extracting specialised features in the pre-processing step. Unlike recent trends in literature, which often focuses on designing frontends or filters for only the initial layers of CNNs, our research introduces the Cosine Convolutional Neural Network (CosCovNN) replacing the traditional CNN filters with Cosine filters. The CosCovNN surpasses the accuracy of the equivalent CNN architectures with approximately $77\%$ less parameters. Our research further progresses with the development of an augmented CosCovNN named Vector Quantised Cosine Convolutional Neural Network with Memory (VQCCM), incorporating a memory and vector quantisation layer VQCCM achieves state-of-the-art (SOTA) performance across five different datasets in comparison with existing literature. Our findings show that cosine filters can greatly improve the efficiency and accuracy of CNNs in raw audio classification.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio ClassificationSimilar Papers 제목 키워드 기반
A Multi-Head Relevance Weighting Framework For Learning Raw Waveform Audio Representations
In this work, we propose a multi-head relevance weighting framework to learn audio representations from raw waveforms. The audio waveform, split into windows of short duration, are processed with a 1-D convolutional laye…
Audio ClassificationSound ClassificationGeneralized zero-shot audio-to-intent classification
Spoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generalized zero-shot audio-to-intent classifica…
ClassificationGoal-Oriented Dialogintent-classificationIntent Classification+4Adapting Language-Audio Models as Few-Shot Audio Learners
We presented the Treff adapter, a training-efficient adapter for CLAP, to boost zero-shot classification performance by making use of a small set of labelled data. Specifically, we designed CALM to retrieve the probabili…
Audio ClassificationClassificationFew-Shot Learningzero-shot-classification+1MP3net: coherent, minute-long music generation from raw audio with a simple convolutional GAN
We present a deep convolutional GAN which leverages techniques from MP3/Vorbis audio compression to produce long, high-quality audio samples with long-range coherence. The model uses a Modified Discrete Cosine Transform …
Audio CompressionMusic GenerationFoolHD: Fooling speaker identification by Highly imperceptible adversarial Disturbances
Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial a…
Adversarial AttackSpeaker Identification