paper-with-me

Papers

Spatio-Temporal Attention Pooling for Audio Scene Classification

2019-04-06 · Huy Phan, Oliver Y. Chén, Lam Pham, Philipp Koch, Maarten De Vos, Ian McLoughlin, Alfred Mertins

Acoustic scenes are rich and redundant in their content. In this work, we present a spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network to learn from patterns that are discriminative while suppressing those that are irrelevant for acoustic scene classification. The convolutional layers in this network learn invariant features from time-frequency input. The bidirectional recurrent layers are then able to encode the temporal dynamics of the resulting convolutional features. Afterwards, a two-dimensional attention mask is formed via the outer product of the spatial and temporal attention vectors learned from two designated attention layers to weigh and pool the recurrent output into a final feature vector for classification. The network is trained with between-class examples generated from between-class data augmentation. Experiments demonstrate that the proposed method not only outperforms a strong convolutional neural network baseline but also sets new state-of-the-art performance on the LITIS Rouen dataset.

📄 PDF Abstract BibTeX arXiv:1904.03543

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic Scene ClassificationClassificationData AugmentationGeneral ClassificationScene Classification

Similar Papers 제목 키워드 기반

Space-Time Memory Network for Sounding Object Localization in Videos

2021-11-10 · Sizhe Li, Yapeng Tian, Chenliang Xu

Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object loc…

Object Localization

Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation

2024-06-10 · Juhyeong Seon, Woobin Im, Sebin Lee, Jumin Lee 외

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensi…

Decoder

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

2026-08-26 · Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa 외 arxiv

We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target …

Spatio-Temporal Action LocalizationContrastive Learning

Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading

2019-10-01 · ICCV 2019 10 · Xingxuan Zhang, Feng Cheng, Shilin Wang

Current state-of-the-art approaches for lip reading are based on sequence-to-sequence architectures that are designed for natural machine translation and audio speech recognition. Hence, these methods do not fully exploi…

LipreadingLip ReadingMachine Translationspeech-recognition+2

Learning to Answer Questions in Dynamic Audio-Visual Scenarios

2022-03-26 · CVPR 2022 1 · Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 외

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive …

audio-visual learningAudio-visual Question AnsweringAudio-Visual Question Answering (AVQA)AUDIO-VISUAL QUESTION ANSWERING (MUSIC-AVQA-v2.0)+4