Papers Environmental Sound Classification
“Environmental Sound Classification” 태그가 달린 논문 50편 · 필터 해제
MAEB: Massive Audio Embedding Benchmark
We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 50+ mod…
Environmental Sound ClassificationExpressive Range Characterization of Open Text-to-Audio Models
Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playabi…
Environmental Sound ClassificationCompressing Quaternion Convolutional Neural Networks for Audio Classification
Conventional Convolutional Neural Networks (CNNs) in the real domain have been widely used for audio classification. However, their convolution operations process multi-channel inputs independently, limiting the ability …
Environmental Sound ClassificationSpeech Emotion RecognitionMusic Genre RecognitionKnowledge DistillationASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
In recent advancements in audio self-supervised representation learning, the standard Transformer architecture has emerged as the predominant approach, yet its attention mechanism often allocates a portion of attention w…
Environmental Sound ClassificationRepresentation LearningAudio ClassificationKeyword SpottingDomain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification
Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely drops in the presence of background sound …
ClassificationDomain AdaptationEnvironmental Sound ClassificationSound ClassificationWeakly Supervised Convolutional Dictionary Learning for Multi-Label Classification
Convolutional Dictionary Learning (CDL) has emerged as a powerful approach for signal representation by learning translation-invariant features through convolution operations. While existing CDL methods are predominantly…
ClassificationDictionary LearningEnvironmental Sound ClassificationMulti-Label Classification+2Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed t…
Audio ClassificationAudio TaggingClassificationEnvironmental Sound Classification+9ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
Environment Sound Classification has been a well-studied research problem in the field of signal processing and up till now more focus has been laid on fully supervised approaches. Over the last few years, focus has move…
Contrastive LearningEnvironmental Sound ClassificationEnvironment Sound ClassificationLanguage Modeling+3Studying the Effect of Audio Filters in Pre-Trained Models for Environmental Sound Classification
Environmental Sound Classification is an important problem of sound recognition and is more complicated than speech recognition problems as environmental sounds are not well structured with respect to time and frequency.…
ClassificationEnvironmental Sound ClassificationSound Classificationspeech-recognition+1Synthetic training set generation using text-to-audio models for environmental sound classification
In recent years, text-to-audio models have revolutionized the field of automatic audio generation. This paper investigates their application in generating synthetic datasets for training data-driven models. Specifically,…
Audio GenerationClassificationEnvironmental Sound ClassificationSound ClassificationMixer is more than just a model
Recently, MLP structures have regained popularity, with MLP-Mixer standing out as a prominent example. In the field of computer vision, MLP-Mixer is noted for its ability to extract data information from both channel and…
Audio ClassificationEnvironmental Sound ClassificationmodelSpeech Emotion RecognitionFocal Modulation Networks for Interpretable Sound Classification
The increasing success of deep neural networks has raised concerns about their inherent black-box nature, posing challenges related to interpretability and trust. While there has been extensive exploration of interpretat…
ClassificationEnvironmental Sound ClassificationSound ClassificationEchoVest: Real-Time Sound Classification and Depth Perception Expressed through Transcutaneous Electrical Nerve Stimulation
Over 1.5 billion people worldwide live with hearing impairment. Despite various technologies that have been created for individuals with such disabilities, most of these technologies are either extremely expensive or ina…
blind source separationClassificationEnvironmental Sound ClassificationSound ClassificationFace: Fast, Accurate and Context-Aware Audio Annotation and Classification
This paper presents a context-aware framework for feature selection and classification procedures to realize a fast and accurate audio event annotation and classification. The context-aware design starts with exploring f…
Active LearningAudio ClassificationClassificationEnvironmental Sound Classification+3Effective Audio Classification Network Based on Paired Inverse Pyramid Structure and Dense MLP Block
Recently, massive architectures based on Convolutional Neural Network (CNN) and self-attention mechanisms have become necessary for audio classification. While these techniques are state-of-the-art, these works' effectiv…
Audio ClassificationClassificationData AugmentationEnvironmental Sound Classification+3Audio Barlow Twins: Self-Supervised Audio Representation Learning
The Barlow Twins self-supervised learning objective requires neither negative samples or asymmetric learning updates, achieving results on a par with the current state-of-the-art within Computer Vision. As such, we prese…
Environmental Sound ClassificationEvent DetectionRepresentation LearningSelf-Supervised LearningImproved Zero-Shot Audio Tagging & Classification with Patchout Spectrogram Transformers
Standard machine learning models for tagging and classifying acoustic signals cannot handle classes that were not seen during training. Zero-Shot (ZS) learning overcomes this restriction by predicting classes based on ad…
Audio TaggingClassificationEnvironmental Sound ClassificationSound ClassificationContinual Learning For On-Device Environmental Sound Classification
Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device environmental sound classification given the restrictions on computation resources (e.g., model size, running memor…
ClassificationComputational EfficiencyContinual LearningEnvironmental Sound Classification+1PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit
PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code struct…
AllAutomatic Speech Recognition (ASR)Environmental Sound ClassificationKeyword Spotting+12End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network
While efficient architectures and a plethora of augmentations for end-to-end image classification tasks have been suggested and heavily investigated, state-of-the-art techniques for audio classifications still rely on nu…
Audio ClassificationClassificationEnvironmental Sound Classificationimage-classification+3