A Hybrid System of Sound Event Detection Transformer and Frame-wise Model for DCASE 2022 Task 4
In this paper, we describe in detail our system for DCASE 2022 Task4. The system combines two considerably different models: an end-to-end Sound Event Detection Transformer (SEDT) and a frame-wise model, Metric Learning and Focal Loss CNN (MLFL-CNN). The former is an event-wise model which learns event-level representations and predicts sound event categories and boundaries directly, while the latter is based on the widely adopted frame-classification scheme, under which each frame is classified into event categories and event boundaries are obtained by post-processing such as thresholding and smoothing. For SEDT, self-supervised pre-training using unlabeled data is applied, and semi-supervised learning is adopted by using an online teacher, which is updated from the student model using the Exponential Moving Average (EMA) strategy and generates reliable pseudo labels for weakly-labeled and unlabeled data. For the frame-wise model, the ICT-TOSHIBA system of DCASE 2021 Task 4 is used. Experimental results show that the hybrid system considerably outperforms either individual model and achieves psds1 of 0.420 and psds2 of 0.783 on the validation set without external data. The code is available at https://github.com/965694547/Hybrid-system-of-frame-wise-model-and-SEDT.
Code (1)
Tasks
Event DetectionMetric LearningSound Event DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Sound Event Detection Transformer: An Event-based End-to-End Model for Sound Event Detection
Sound event detection (SED) has gained increasing attention with its wide application in surveillance, video indexing, etc. Existing models in SED mainly generate frame-level prediction, converting it into a sequence mul…
Audio TaggingBoundary DetectionEvent DetectionMulti-Label Classification+5Effective Pre-Training of Audio Transformers for Sound Event Detection
We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-le…
Data AugmentationEvent DetectionKnowledge DistillationSound Event DetectionA hybrid parametric-deep learning approach for sound event localization and detection
This work describes and discusses an algorithm submitted to the Sound Event Localization and Detection Task of DCASE2019 Challenge. The proposed methodology relies on parametric spatial audio analysis for source localiza…
Sound Event Localization and DetectionAudioLog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning
Previous studies in automated audio captioning have faced difficulties in accurately capturing the complete temporal details of acoustic scenes and events within long audio sequences. This paper presents AudioLog, a larg…
Acoustic Scene ClassificationAudio captioningContrastive LearningEvent Detection+3PILOT: Introducing Transformers for Probabilistic Sound Event Localization
Sound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in this domain most prominently focused on uti…
Event Detection