Spatio-Temporal Attention Pooling for Audio Scene Classification
Acoustic scenes are rich and redundant in their content. In this work, we present a spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network to learn from patterns that are discriminative while suppressing those that are irrelevant for acoustic scene classification. The convolutional layers in this network learn invariant features from time-frequency input. The bidirectional recurrent layers are then able to encode the temporal dynamics of the resulting convolutional features. Afterwards, a two-dimensional attention mask is formed via the outer product of the spatial and temporal attention vectors learned from two designated attention layers to weigh and pool the recurrent output into a final feature vector for classification. The network is trained with between-class examples generated from between-class data augmentation. Experiments demonstrate that the proposed method not only outperforms a strong convolutional neural network baseline but also sets new state-of-the-art performance on the LITIS Rouen dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Acoustic Scene ClassificationClassificationData AugmentationGeneral ClassificationScene ClassificationSimilar Papers 제목 키워드 기반
Space-Time Memory Network for Sounding Object Localization in Videos
Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object loc…
Object LocalizationExtending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensi…
DecoderSkeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining
We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target …
Spatio-Temporal Action LocalizationContrastive LearningSpatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading
Current state-of-the-art approaches for lip reading are based on sequence-to-sequence architectures that are designed for natural machine translation and audio speech recognition. Hence, these methods do not fully exploi…
LipreadingLip ReadingMachine Translationspeech-recognition+2Learning to Answer Questions in Dynamic Audio-Visual Scenarios
In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive …
audio-visual learningAudio-visual Question AnsweringAudio-Visual Question Answering (AVQA)AUDIO-VISUAL QUESTION ANSWERING (MUSIC-AVQA-v2.0)+4