paper-with-me

Audio Tagging

1개 벤치마크 · 논문 85편 · 이 태스크의 논문 보기 →

Benchmarks

AudioSet

결과 11개

Most implemented

AST: Audio Spectrogram Transformer

2021-04-05 · 구현 5개

Speech Denoising with Deep Feature Losses

2018-06-27 · 구현 5개

Papers

From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

2026-06-26 · Stefano Giacomelli, Stefano Damiano, Claudia Rinaldi, Fabio Graziosi 외 arxiv

This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrain…

Sound Event Localization and DetectionNeural Architecture SearchAudio Tagging

Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context

2026-03-11 · Yuanbo Hou, Yanru Wu, Qiaoqiao Ren, Shengchen Li 외 arxiv

Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT):…

Audio Tagging

Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices

2025-09-17 · Jordi Grau-Haro, Ruben Ribes-Serrano, Javier Naranjo-Alcazar, Marta Garcia-Ballesteros 외 arxiv

Convolutional Neural Networks (CNNs) have demonstrated exceptional performance in audio tagging tasks. However, deploying these models on resource-constrained devices like the Raspberry Pi poses challenges related to com…

Computational EfficiencyAudio ClassificationAudio Tagging

On Temporal Guidance and Iterative Refinement in Audio Source Separation

2025-07-23 · Tobias Morocutti, Jonathan Greif, Paul Primus, Florian Schmid 외 arxiv

Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a t…

Audio Source SeparationSemantic SegmentationSound Event DetectionAudio Tagging

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4

2025-06-26 · Jongyeon Park, Joonhee Lee, Do-Hyeon Lim, Hong Kook Kim 외

This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted fro…

Audio TaggingSemantic Segmentation

USAD: Universal Speech and Audio Representation via Distillation

2025-06-23 · Heng-Jui Chang, Saurabhchand Bhati, James Glass, Alexander H. Liu

Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distill…

Audio TaggingRepresentation LearningSelf-Supervised LearningSound Classification

전체 85편 보기 →