Unsupervised Discriminative Learning of Sounds for Audio Event Classification
Recent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale visual datasets is time consuming. On several audio event classification benchmarks, we show a fast and effective alternative that pre-trains the model unsupervised, only on audio data and yet delivers on-par performance with ImageNet pre-training. Furthermore, we show that our discriminative audio learning can be used to transfer knowledge across audio datasets and optionally include ImageNet pre-training.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationTransfer LearningSimilar Papers 제목 키워드 기반
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
Open-vocabulary audio-language models, like CLAP, offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language promp…
Audio ClassificationDescriptiveText RetrievalZero-shot Audio ClassificationPlay It Back: Iterative Attention for Audio Recognition
A key function of auditory cognition is the association of characteristic sounds with their corresponding semantics over time. Humans attempting to discriminate between fine-grained audio categories, often replay the sam…
Audio ClassificationA dataset for Audio-Visual Sound Event Detection in Movies
Audio event detection is a widely studied audio processing task, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many …
Event DetectionSelf-Driving CarsSound ClassificationSound Event DetectionUnsupervised Learning of Semantic Audio Representations
Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds. We consider several class-agnostic semantic con…
Audio ClassificationClassificationGeneral ClassificationRetrieval+1Learning domain-invariant classifiers for infant cry sounds
The issue of domain shift remains a problematic phenomenon in most real-world datasets and clinical audio is no exception. In this work, we study the nature of domain shift in a clinical database of infant cry sounds acq…
Domain AdaptationUnsupervised Domain Adaptation