Synthetic data enables context-aware bioacoustic sound event detection
We propose a methodology for training foundation models that enhances their in-context learning capabilities within the domain of bioacoustic signal processing. We use synthetically generated training data, introducing a domain-randomization-based pipeline that constructs diverse acoustic scenes with temporally strong labels. We generate over 8.8 thousand hours of strongly-labeled audio and train a query-by-example, transformer-based model to perform few-shot bioacoustic sound event detection. Our second contribution is a public benchmark of 13 diverse few-shot bioacoustics tasks. Our model outperforms previously published methods by 49%, and we demonstrate that this is due to both model design and data scale. We make our trained model available via an API, to provide ecologists and ethologists with a training-free tool for bioacoustic sound event detection.
Code (0)
등록된 구현이 없습니다.
Tasks
Event DetectionIn-Context LearningSound Event DetectionSimilar Papers 제목 키워드 기반
Towards Deep Active Learning in Avian Bioacoustics
Passive acoustic monitoring (PAM) in avian bioacoustics enables cost-effective and extensive data collection with minimal disruption to natural habitats. Despite advancements in computational avian bioacoustics, deep lea…
Active LearningThe bioacoustic proof of the effects of raising awareness of noise pollution among visitors to the Port Cros National Park using binding communication
Assuming that the anthropogenic impact of visitors to a natural park can be reduced by communication actions, we have drawn up a bioacoustical protocol combined with a protocol to measure the effectiveness of the communi…
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
Large language models (LLMs) prompted with text and audio represent the state of the art in various auditory tasks, including speech, music, and general audio, showing emergent abilities on unseen tasks. However, these c…
zero-shot-classificationZero-Shot LearningKnowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis
Marine mammal vocalization analysis depends on interpreting bioacoustic spectrograms. Vision Language Models (VLMs) are not trained on these domain-specific visualizations. We investigate whether VLMs can extract meaning…
Multi-layer attentive probing improves transfer of audio representations for bioacoustics
Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating representation learning. Most bioacoustic benchmarks use a fixed, low-c…
Representation Learning