CNNs-based Acoustic Scene Classification using Multi-Spectrogram Fusion and Label Expansions
Spectrograms have been widely used in Convolutional Neural Networks based schemes for acoustic scene classification, such as the STFT spectrogram and the MFCC spectrogram, etc. They have different time-frequency characteristics, contributing to their own advantages and disadvantages in recognizing acoustic scenes. In this letter, a novel multi-spectrogram fusion framework is proposed, making the spectrograms complement each other. In the framework, a single CNN architecture is applied onto multiple spectrograms for feature extraction. The deep features extracted from multiple spectrograms are then fused to discriminate the acoustic scenes. Moreover, motivated by the inter-class similarities in acoustic scene datasets, a label expansion method is further proposed in which super-class labels are constructed upon the original classes. On the help of the expanded labels, the CNN models are transformed into the multitask learning form to improve the acoustic scene classification by appending the auxiliary task of super-class classification. To verify the effectiveness of the proposed methods, intensive experiments have been performed on the DCASE2017 and the LITIS Rouen datasets. Experimental results show that the proposed method can achieve promising accuracies on both datasets. Specifically, accuracies of 0.9744, 0.8865 and 0.7778 are obtained for the LITIS Rouen dataset, the DCASE Development set and Evaluation set respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
Acoustic Scene ClassificationClassificationGeneral ClassificationScene ClassificationSimilar Papers 제목 키워드 기반
A Comparative Study on Approaches to Acoustic Scene Classification using CNNs
Acoustic scene classification is a process of characterizing and classifying the environments from sound recordings. The first step is to generate features (representations) from the recorded sound and then classify the …
Acoustic Scene ClassificationClassificationScene ClassificationSpectNet : End-to-End Audio Signal Classification Using Learnable Spectrograms
Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MF…
Acoustic Scene ClassificationAnomaly DetectionAudio ClassificationAudio Tagging+4Acoustic Scene Classification Using Fusion of Attentive Convolutional Neural Networks for DCASE2019 Challenge
In this report, the Brno University of Technology (BUT) team submissions for Task 1 (Acoustic Scene Classification, ASC) of the DCASE-2019 challenge are described. Also, the analysis of different methods is provided. The…
Acoustic Scene ClassificationScene ClassificationSpeaker VerificationLow-complexity deep learning frameworks for acoustic scene classification
In this report, we presents low-complexity deep learning frameworks for acoustic scene classification (ASC). The proposed frameworks can be separated into four main steps: Front-end spectrogram extraction, online data au…
Acoustic Scene ClassificationClassificationData AugmentationDeep Learning+1Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
The ability to generalize to a wide range of recording devices is a crucial performance factor for audio classification models. The characteristics of different types of microphones introduce distributional shifts in the…
Acoustic Scene ClassificationAudio ClassificationClassificationDiversity+1