paper-with-me

홈 › Papers

Audiovisual transfer learning for audio tagging and sound event detection

2021-06-09 · Wim Boes, Hugo Van hamme

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of pretrained auditory and visual features, extracted from networks built for different tasks and trained with external data. We perform experiments with these modified models on an audiovisual multi-label data set, of which the training partition contains a large number of unlabeled samples and a smaller amount of clips with weak annotations, indicating the clip-level presence of 10 sound categories without specifying the temporal boundaries of the active auditory events. For clip-based audio tagging, this transfer learning method grants marked improvements. Addition of the visual modality on top of audio also proves to be advantageous in this context. When it comes to generating transcriptions of audio recordings, the benefit of pretrained features depends on the requested temporal resolution: for coarse-grained sound event detection, their utility remains notable. But when more fine-grained predictions are required, performance gains are strongly reduced due to a mismatch between the problem at hand and the goals of the models from which the pretrained vectors were obtained.

📄 PDF Abstract BibTeX arXiv:2106.05408

Code (0)

등록된 구현이 없습니다.

Tasks

Audio TaggingEvent DetectionSound Event DetectionTransfer Learning

Similar Papers 제목 키워드 기반

Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition

2020-05-18 · ECCV 2020 8 · Di Hu, Xuhong LI, Lichao Mou, Pu Jin 외

Aerial scene recognition is a fundamental task in remote sensing and has recently received increased interest. While the visual information from overhead images with powerful models and efficient algorithms yields consid…

Scene Recognition

Class-aware Sounding Objects Localization via Audiovisual Correspondence

2021-12-22 · Di Hu, Yake Wei, Rui Qian, Weiyao Lin 외

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localiza…

Objectobject-detectionObject DetectionObject Localization+1

Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data

2020-05-29 · Haytham M. Fayek, Anurag Kumar

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to diff…

Audio Classification

Audio Tagging With Connectionist Temporal Classification Model Using Sequential Labelled Data

2018-08-06 · Yuanbo Hou, Qiuqiang Kong, Shengchen Li

Audio tagging aims to predict one or several labels in an audio clip. Many previous works use weakly labelled data (WLD) for audio tagging, where only presence or absence of sound events is known, but the order of sound …

Audio TaggingGeneral Classification

Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation

2025-09-26 · Jinbae Seo, Hyeongjun Kwon, Kwonyoung Kim, Jiyoung Lee 외 arxiv

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform add…

Instance Segmentation