paper-with-me

홈 › Papers

Simple Pooling Front-ends For Efficient Audio Classification

2022-10-03 · Xubo Liu, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Mark D. Plumbley, Wenwu Wang

Recently, there has been increasing interest in building efficient audio neural networks for on-device scenarios. Most existing approaches are designed to reduce the size of audio neural networks using methods such as model pruning. In this work, we show that instead of reducing model size using complex methods, eliminating the temporal redundancy in the input audio features (e.g., mel-spectrogram) could be an effective approach for efficient audio classification. To do so, we proposed a family of simple pooling front-ends (SimPFs) which use simple non-parametric pooling operations to reduce the redundant information within the mel-spectrogram. We perform extensive experiments on four audio classification tasks to evaluate the performance of SimPFs. Experimental results show that SimPFs can achieve a reduction in more than half of the number of floating point operations (FLOPs) for off-the-shelf audio neural networks, with negligible degradation or even some improvements in audio classification performance.

📄 PDF Abstract BibTeX arXiv:2210.00943

Code (1)

liuxubo717/simpfs 공식 구현 pytorch

Tasks

Audio ClassificationClassification

Similar Papers 제목 키워드 기반

LEAF: A Learnable Frontend for Audio Classification

2021-01-21 · Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco Tagliasacchi

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by th…

Audio ClassificationClassificationGeneral Classification

Analysis of ABC Frontend Audio Systems for the NIST-SRE24

2025-05-21 · Sara Barahona, Anna Silnova, Ladislav Mošner, Junyi Peng 외

We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of teleph…

Speaker Recognition

Robust Audio Event Recognition with 1-Max Pooling Convolutional Neural Networks

2016-04-21 · Huy Phan, Lars Hertel, Marco Maass, Alfred Mertins

We present in this paper a simple, yet efficient convolutional neural network (CNN) architecture for robust audio event recognition. Opposing to deep CNN architectures with multiple convolutional and pooling layers toppe…

A Universal Learnable Audio Frontend

2021-01-01 · ICLR 2021 1 · Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco Tagliasacchi

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have lived through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fu…

Audio Classification

EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use

2022-07-12 · Jan Schlüter, Gerald Gutenbrunner

In audio classification, differentiable auditory filterbanks with few parameters cover the middle ground between hard-coded spectrograms and raw audio. LEAF (arXiv:2101.08596), a Gabor-based filterbank combined with Per-…

Audio ClassificationClassificationInstrument RecognitionPitch Classification+1