paper-with-me

Papers Audio Tagging

“Audio Tagging” 태그가 달린 논문 85편 · 필터 해제

From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

2026-06-26 · Stefano Giacomelli, Stefano Damiano, Claudia Rinaldi, Fabio Graziosi 외 arxiv

This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrain…

Sound Event Localization and DetectionNeural Architecture SearchAudio Tagging

Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context

2026-03-11 · Yuanbo Hou, Yanru Wu, Qiaoqiao Ren, Shengchen Li 외 arxiv

Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT):…

Audio Tagging

Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices

2025-09-17 · Jordi Grau-Haro, Ruben Ribes-Serrano, Javier Naranjo-Alcazar, Marta Garcia-Ballesteros 외 arxiv

Convolutional Neural Networks (CNNs) have demonstrated exceptional performance in audio tagging tasks. However, deploying these models on resource-constrained devices like the Raspberry Pi poses challenges related to com…

Computational EfficiencyAudio ClassificationAudio Tagging

On Temporal Guidance and Iterative Refinement in Audio Source Separation

2025-07-23 · Tobias Morocutti, Jonathan Greif, Paul Primus, Florian Schmid 외 arxiv

Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a t…

Audio Source SeparationSemantic SegmentationSound Event DetectionAudio Tagging

Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4

2025-06-26 · Jongyeon Park, Joonhee Lee, Do-Hyeon Lim, Hong Kook Kim 외

This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted fro…

Audio TaggingSemantic Segmentation

USAD: Universal Speech and Audio Representation via Distillation

2025-06-23 · Heng-Jui Chang, Saurabhchand Bhati, James Glass, Alexander H. Liu

Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distill…

Audio TaggingRepresentation LearningSelf-Supervised LearningSound Classification

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

2025-06-02 · Alef Iury Siqueira Ferreira, Lucas Rafael Gris, Alexandre Ferro Filho, Lucas Ólives 외

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the IN…

Audio TaggingEmotion RecognitionGraph AttentionQuantization+1

Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes

2025-03-28 · Binh Thien Nguyen, Masahiro Yasuda, Daiki Takeuchi, Daisuke Niizumi 외

Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introdu…

Audio TaggingSemantic Segmentation

M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

2025-03-28 · Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen 외

Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …

Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13

Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging

2025-03-26 · Ludovic Tuncay, Etienne Labbé, Thomas Pellegrini

AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology. However, the annotations co…

Audio Tagging

Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context

2025-03-19 · Junyi Ao, Dekun Chen, Xiaohai Tian, Wenjie Feng 외

Large Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio. However, most existing models primarily focus on analyzing input signals u…

Audio captioningAudio Question AnsweringAudio TaggingQuestion Answering

Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models

2025-03-14 · Tobias Morocutti, Florian Schmid, Jonathan Greif, Francesco Foscarin 외

We target the problem of developing new low-complexity networks for the sound event detection task. Our goal is to meticulously analyze the performance-complexity trade-off, aiming to be competitive with the large state-…

Audio TaggingEvent DetectionKnowledge DistillationSound Event Detection

Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning

2025-02-17 · ICASSP 2025 3 · Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid

Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed t…

Audio ClassificationAudio TaggingClassificationEnvironmental Sound Classification+9

Knowledge Distillation for Real-Time Classification of Early Media in Voice Communications

2024-10-28 · Kemal Altwlkany, Hadžem Hadžić, Amar Kurić, Emanuel Lacic

This paper investigates the industrial setting of real-time classification of early media exchanged during the initialization phase of voice calls. We explore the application of state-of-the-art audio tagging models and …

Audio TaggingClassificationKnowledge Distillation

MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events

2024-09-25 · Xiaoyu Yang, Qiujia Li, Chao Zhang, Phil Woodland

With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still challenging to build a general-purpose model …

Audio TaggingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillation+4

LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging

2024-09-17 · Charilaos Papaioannou, Emmanouil Benetos, Alexandros Potamianos

We introduce Label-Combination Prototypical Networks (LC-Protonets) to address the problem of multi-label few-shot classification, where a model must generalize to new classes based on only a few available examples. Exte…

Audio TaggingFew-Shot LearningMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Normalizing Energy Consumption for Hardware-Independent Evaluation

2024-09-09 · Constance Douwes, Romain Serizel

The increasing use of machine learning (ML) models in signal processing has raised concerns about their environmental impact, particularly during resource-intensive training phases. In this study, we present a novel meth…

Audio Tagging

From Computation to Consumption: Exploring the Compute-Energy Link for Training and Testing Neural Networks for SED Systems

2024-09-08 · Constance Douwes, Romain Serizel

The massive use of machine learning models, particularly neural networks, has raised serious concerns about their environmental impact. Indeed, over the last few years we have seen an explosion in the computing costs ass…

Audio TaggingEvent DetectionGPUSound Event Detection

A Framework for Synthetic Audio Conversations Generation using Large Language Models

2024-09-02 · Kaung Myat Kyaw, Jonathan Hoyin Chan

In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent…

Audio ClassificationAudio TaggingDiversityspeech-recognition+3

Integrating IP Broadcasting with Audio Tags: Workflow and Challenges

2024-07-22 · Rhys Burchett-Vass, Arshdeep Singh, Gabriel Bibbó, Mark D. Plumbley

The broadcasting industry is increasingly adopting IP techniques, revolutionising both live and pre-recorded content production, from news gathering to live music events. IP broadcasting allows for the transport of audio…

Audio Tagging
1–20 / 85 다음 →