AclNet: efficient end-to-end audio classification CNN
We propose an efficient end-to-end convolutional neural network architecture, AclNet, for audio classification. When trained with our data augmentation and regularization, we achieved state-of-the-art performance on the ESC-50 corpus with 85:65% accuracy. Our network allows configurations such that memory and compute requirements are drastically reduced, and a tradeoff analysis of accuracy and complexity is presented. The analysis shows high accuracy at significantly reduced computational complexity compared to existing solutions. For example, a configuration with only 155k parameters and 49:3 million multiply-adds per second is 81:75%, exceeding human accuracy of 81:3%. This improved efficiency can enable always-on inference in energy-efficient platforms.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio ClassificationClassificationData AugmentationGeneral ClassificationSimilar Papers 제목 키워드 기반
ACLNet: An Attention and Clustering-based Cloud Segmentation Network
We propose a novel deep learning model named ACLNet, for cloud segmentation from ground images. ACLNet uses both deep neural network and machine learning (ML) algorithm to extract complementary features. Specifically, it…
ClusteringSegmentationSemantic SegmentationContext, Attention and Audio Feature Explorations for Audio Visual Scene-Aware Dialog
With the recent advancements in AI, Intelligent Virtual Assistants (IVA) have become a ubiquitous part of every home. Going forward, we are witnessing a confluence of vision, speech and dialog system technologies that ar…
Audio ClassificationGeneral ClassificationLeveraging Topics and Audio Features with Multimodal Attention for Audio Visual Scene-Aware Dialog
With the recent advancements in Artificial Intelligence (AI), Intelligent Virtual Assistants (IVA) such as Alexa, Google Home, etc., have become a ubiquitous part of many homes. Currently, such IVAs are mostly audio-base…
Audio ClassificationResponse GenerationExploring Context, Attention and Audio Features for Audio Visual Scene-Aware Dialog
We are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs to learn audio-visual groundings of utterances and have conversations with users about the objects, activities an…
Audio ClassificationVisual GroundingAffinity Contrastive Learning for Skeleton-based Human Activity Understanding
In skeleton-based human activity understanding, existing methods often adopt the contrastive learning paradigm to construct a discriminative feature space. However, many of these approaches fail to exploit the structural…
Person Re-IdentificationContrastive LearningAction RecognitionGait Recognition