paper-with-me

Papers

Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)

2024-11-30 · Kazi Nazmul Haque, Rajib Rana, Tasnim Jarin, Bjorn W. Schuller Jr

This study explores the field of audio classification from raw waveform using Convolutional Neural Networks (CNNs), a method that eliminates the need for extracting specialised features in the pre-processing step. Unlike recent trends in literature, which often focuses on designing frontends or filters for only the initial layers of CNNs, our research introduces the Cosine Convolutional Neural Network (CosCovNN) replacing the traditional CNN filters with Cosine filters. The CosCovNN surpasses the accuracy of the equivalent CNN architectures with approximately $77\%$ less parameters. Our research further progresses with the development of an augmented CosCovNN named Vector Quantised Cosine Convolutional Neural Network with Memory (VQCCM), incorporating a memory and vector quantisation layer VQCCM achieves state-of-the-art (SOTA) performance across five different datasets in comparison with existing literature. Our findings show that cosine filters can greatly improve the efficiency and accuracy of CNNs in raw audio classification.

📄 PDF Abstract BibTeX arXiv:2412.00312

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Classification

Similar Papers 제목 키워드 기반

A Multi-Head Relevance Weighting Framework For Learning Raw Waveform Audio Representations

2021-07-30 · Debottam Dutta, Purvi Agrawal, Sriram Ganapathy

In this work, we propose a multi-head relevance weighting framework to learn audio representations from raw waveforms. The audio waveform, split into windows of short duration, are processed with a 1-D convolutional laye…

Audio ClassificationSound Classification

Generalized zero-shot audio-to-intent classification

2023-11-04 · Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi, Sravan Bodapati 외

Spoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generalized zero-shot audio-to-intent classifica…

ClassificationGoal-Oriented Dialogintent-classificationIntent Classification+4

Adapting Language-Audio Models as Few-Shot Audio Learners

2023-05-28 · Jinhua Liang, Xubo Liu, Haohe Liu, Huy Phan 외

We presented the Treff adapter, a training-efficient adapter for CLAP, to boost zero-shot classification performance by making use of a small set of labelled data. Specifically, we designed CALM to retrieve the probabili…

Audio ClassificationClassificationFew-Shot Learningzero-shot-classification+1

MP3net: coherent, minute-long music generation from raw audio with a simple convolutional GAN

2021-01-12 · Korneel van den Broek

We present a deep convolutional GAN which leverages techniques from MP3/Vorbis audio compression to produce long, high-quality audio samples with long-range coherence. The model uses a Modified Discrete Cosine Transform …

Audio CompressionMusic Generation

FoolHD: Fooling speaker identification by Highly imperceptible adversarial Disturbances

2020-11-17 · Ali Shahin Shamsabadi, Francisco Sepúlveda Teixeira, Alberto Abad, Bhiksha Raj 외

Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial a…

Adversarial AttackSpeaker Identification