Audio Tagging
1개 벤치마크 · 논문 85편 · 이 태스크의 논문 보기 →
Benchmarks
AudioSet
Most implemented
AST: Audio Spectrogram Transformer
Speech Denoising with Deep Feature Losses
musicnn: Pre-trained convolutional neural networks for music audio tagging
General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
Audio classification with Dilated Convolution with Learnable Spacings
Papers
From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection
This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrain…
Sound Event Localization and DetectionNeural Architecture SearchAudio TaggingGeo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT):…
Audio TaggingComprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
Convolutional Neural Networks (CNNs) have demonstrated exceptional performance in audio tagging tasks. However, deploying these models on resource-constrained devices like the Raspberry Pi poses challenges related to com…
Computational EfficiencyAudio ClassificationAudio TaggingOn Temporal Guidance and Iterative Refinement in Audio Source Separation
Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a t…
Audio Source SeparationSemantic SegmentationSound Event DetectionAudio TaggingPerformance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4
This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted fro…
Audio TaggingSemantic SegmentationUSAD: Universal Speech and Audio Representation via Distillation
Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distill…
Audio TaggingRepresentation LearningSelf-Supervised LearningSound Classification