paper-with-me

Papers Environmental Sound Classification

“Environmental Sound Classification” 태그가 달린 논문 50편 · 필터 해제

MAEB: Massive Audio Embedding Benchmark

2026-02-17 · Adnan El Assadi, Isaac Chung, Chenghao Xiao, Roman Solomatin 외 arxiv

We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 50+ mod…

Environmental Sound Classification

Expressive Range Characterization of Open Text-to-Audio Models

2025-10-31 · Jonathan Morse, Azadeh Naderi, Swen Gaudl, Mark Cartwright 외 arxiv

Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playabi…

Environmental Sound Classification

Compressing Quaternion Convolutional Neural Networks for Audio Classification

2025-10-24 · Arshdeep Singh, Vinayak Abrol, Mark D. Plumbley arxiv

Conventional Convolutional Neural Networks (CNNs) in the real domain have been widely used for audio classification. However, their convolution operations process multi-channel inputs independently, limiting the ability …

Environmental Sound ClassificationSpeech Emotion RecognitionMusic Genre RecognitionKnowledge Distillation

ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning

2025-07-03 · Junyu Wang, Tianrui Wang, Meng Ge, Longbiao Wang 외 arxiv

In recent advancements in audio self-supervised representation learning, the standard Transformer architecture has emerged as the predominant approach, yet its attention mechanism often allocates a portion of attention w…

Environmental Sound ClassificationRepresentation LearningAudio ClassificationKeyword Spotting

Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification

2025-06-04 · Emiliano Acevedo, Martín Rocamora, Magdalena Fuentes

Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely drops in the presence of background sound …

ClassificationDomain AdaptationEnvironmental Sound ClassificationSound Classification

Weakly Supervised Convolutional Dictionary Learning for Multi-Label Classification

2025-03-11 · Hao Chen, Dayuan Tan

Convolutional Dictionary Learning (CDL) has emerged as a powerful approach for signal representation by learning translation-invariant features through convolution operations. While existing CDL methods are predominantly…

ClassificationDictionary LearningEnvironmental Sound ClassificationMulti-Label Classification+2

Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning

2025-02-17 · ICASSP 2025 3 · Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid

Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed t…

Audio ClassificationAudio TaggingClassificationEnvironmental Sound Classification+9

ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning

2024-09-21 · Pranav Gupta, Raunak Sharma, Rashmi Kumari, Sri Krishna Aditya 외

Environment Sound Classification has been a well-studied research problem in the field of signal processing and up till now more focus has been laid on fully supervised approaches. Over the last few years, focus has move…

Contrastive LearningEnvironmental Sound ClassificationEnvironment Sound ClassificationLanguage Modeling+3

Studying the Effect of Audio Filters in Pre-Trained Models for Environmental Sound Classification

2024-08-24 · Aditya Dawn, Wazib Ansar

Environmental Sound Classification is an important problem of sound recognition and is more complicated than speech recognition problems as environmental sounds are not well structured with respect to time and frequency.…

ClassificationEnvironmental Sound ClassificationSound Classificationspeech-recognition+1

Synthetic training set generation using text-to-audio models for environmental sound classification

2024-03-26 · Francesca Ronchini, Luca Comanducci, Fabio Antonacci

In recent years, text-to-audio models have revolutionized the field of automatic audio generation. This paper investigates their application in generating synthetic datasets for training data-driven models. Specifically,…

Audio GenerationClassificationEnvironmental Sound ClassificationSound Classification

Mixer is more than just a model

2024-02-28 · Qingfeng Ji, Yuxin Wang, Letong Sun

Recently, MLP structures have regained popularity, with MLP-Mixer standing out as a prominent example. In the field of computer vision, MLP-Mixer is noted for its ability to extract data information from both channel and…

Audio ClassificationEnvironmental Sound ClassificationmodelSpeech Emotion Recognition

Focal Modulation Networks for Interpretable Sound Classification

2024-02-05 · Luca Della Libera, Cem Subakan, Mirco Ravanelli

The increasing success of deep neural networks has raised concerns about their inherent black-box nature, posing challenges related to interpretability and trust. While there has been extensive exploration of interpretat…

ClassificationEnvironmental Sound ClassificationSound Classification

EchoVest: Real-Time Sound Classification and Depth Perception Expressed through Transcutaneous Electrical Nerve Stimulation

2023-07-10 · Jesse Choe, Siddhant Sood, Ryan Park

Over 1.5 billion people worldwide live with hearing impairment. Despite various technologies that have been created for individuals with such disabilities, most of these technologies are either extremely expensive or ina…

blind source separationClassificationEnvironmental Sound ClassificationSound Classification

Face: Fast, Accurate and Context-Aware Audio Annotation and Classification

2023-03-07 · M. Mehrdad Morsali, Hoda Mohammadzade, Saeed Bagheri Shouraki

This paper presents a context-aware framework for feature selection and classification procedures to realize a fast and accurate audio event annotation and classification. The context-aware design starts with exploring f…

Active LearningAudio ClassificationClassificationEnvironmental Sound Classification+3

Effective Audio Classification Network Based on Paired Inverse Pyramid Structure and Dense MLP Block

2022-11-05 · Yunhao Chen, Yunjie Zhu, Zihui Yan, Yifan Huang 외

Recently, massive architectures based on Convolutional Neural Network (CNN) and self-attention mechanisms have become necessary for audio classification. While these techniques are state-of-the-art, these works' effectiv…

Audio ClassificationClassificationData AugmentationEnvironmental Sound Classification+3

Audio Barlow Twins: Self-Supervised Audio Representation Learning

2022-09-28 · Jonah Anton, Harry Coppock, Pancham Shukla, Bjorn W. Schuller

The Barlow Twins self-supervised learning objective requires neither negative samples or asymmetric learning updates, achieving results on a par with the current state-of-the-art within Computer Vision. As such, we prese…

Environmental Sound ClassificationEvent DetectionRepresentation LearningSelf-Supervised Learning

Improved Zero-Shot Audio Tagging & Classification with Patchout Spectrogram Transformers

2022-08-24 · Paul Primus, Gerhard Widmer

Standard machine learning models for tagging and classifying acoustic signals cannot handle classes that were not seen during training. Zero-Shot (ZS) learning overcomes this restriction by predicting classes based on ad…

Audio TaggingClassificationEnvironmental Sound ClassificationSound Classification

Continual Learning For On-Device Environmental Sound Classification

2022-07-15 · Yang Xiao, Xubo Liu, James King, Arshdeep Singh 외

Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device environmental sound classification given the restrictions on computation resources (e.g., model size, running memor…

ClassificationComputational EfficiencyContinual LearningEnvironmental Sound Classification+1

PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

2022-05-20 · NAACL (ACL) 2022 7 · HUI ZHANG, Tian Yuan, Junkun Chen, Xintong Li 외

PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code struct…

AllAutomatic Speech Recognition (ASR)Environmental Sound ClassificationKeyword Spotting+12

End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network

2022-04-25 · Avi Gazneli, Gadi Zimerman, Tal Ridnik, Gilad Sharir 외

While efficient architectures and a plethora of augmentations for end-to-end image classification tasks have been suggested and heavily investigated, state-of-the-art techniques for audio classifications still rely on nu…

Audio ClassificationClassificationEnvironmental Sound Classificationimage-classification+3
1–20 / 50 다음 →