Papers Audio Tagging
“Audio Tagging” 태그가 달린 논문 85편 · 필터 해제
From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection
This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrain…
Sound Event Localization and DetectionNeural Architecture SearchAudio TaggingGeo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT):…
Audio TaggingComprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
Convolutional Neural Networks (CNNs) have demonstrated exceptional performance in audio tagging tasks. However, deploying these models on resource-constrained devices like the Raspberry Pi poses challenges related to com…
Computational EfficiencyAudio ClassificationAudio TaggingOn Temporal Guidance and Iterative Refinement in Audio Source Separation
Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a t…
Audio Source SeparationSemantic SegmentationSound Event DetectionAudio TaggingPerformance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4
This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted fro…
Audio TaggingSemantic SegmentationUSAD: Universal Speech and Audio Representation via Distillation
Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distill…
Audio TaggingRepresentation LearningSelf-Supervised LearningSound ClassificationEnhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the IN…
Audio TaggingEmotion RecognitionGraph AttentionQuantization+1Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introdu…
Audio TaggingSemantic SegmentationM2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …
Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology. However, the annotations co…
Audio TaggingSolla: Towards a Speech-Oriented LLM That Hears Acoustic Context
Large Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio. However, most existing models primarily focus on analyzing input signals u…
Audio captioningAudio Question AnsweringAudio TaggingQuestion AnsweringExploring Performance-Complexity Trade-Offs in Sound Event Detection Models
We target the problem of developing new low-complexity networks for the sound event detection task. Our goal is to meticulously analyze the performance-complexity trade-off, aiming to be competitive with the large state-…
Audio TaggingEvent DetectionKnowledge DistillationSound Event DetectionMasked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed t…
Audio ClassificationAudio TaggingClassificationEnvironmental Sound Classification+9Knowledge Distillation for Real-Time Classification of Early Media in Voice Communications
This paper investigates the industrial setting of real-time classification of early media exchanged during the initialization phase of voice calls. We explore the application of state-of-the-art audio tagging models and …
Audio TaggingClassificationKnowledge DistillationMT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still challenging to build a general-purpose model …
Audio TaggingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillation+4LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
We introduce Label-Combination Prototypical Networks (LC-Protonets) to address the problem of multi-label few-shot classification, where a model must generalize to new classes based on only a few available examples. Exte…
Audio TaggingFew-Shot LearningMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONNormalizing Energy Consumption for Hardware-Independent Evaluation
The increasing use of machine learning (ML) models in signal processing has raised concerns about their environmental impact, particularly during resource-intensive training phases. In this study, we present a novel meth…
Audio TaggingFrom Computation to Consumption: Exploring the Compute-Energy Link for Training and Testing Neural Networks for SED Systems
The massive use of machine learning models, particularly neural networks, has raised serious concerns about their environmental impact. Indeed, over the last few years we have seen an explosion in the computing costs ass…
Audio TaggingEvent DetectionGPUSound Event DetectionA Framework for Synthetic Audio Conversations Generation using Large Language Models
In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent…
Audio ClassificationAudio TaggingDiversityspeech-recognition+3Integrating IP Broadcasting with Audio Tags: Workflow and Challenges
The broadcasting industry is increasingly adopting IP techniques, revolutionising both live and pre-recorded content production, from news gathering to live music events. IP broadcasting allows for the transport of audio…
Audio Tagging