paper-with-me

Papers

Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics

2024-09-21 · Burooj Ghani, Vincent J. Kalkman, Bob Planqué, Willem-Pier Vellinga, Lisa Gill, Dan Stowell

Animal sounds can be recognised automatically by machine learning, and this has an important role to play in biodiversity monitoring. Yet despite increasingly impressive capabilities, bioacoustic species classifiers still exhibit imbalanced performance across species and habitats, especially in complex soundscapes. In this study, we explore the effectiveness of transfer learning in large-scale bird sound classification across various conditions, including single- and multi-label scenarios, and across different model architectures such as CNNs and Transformers. Our experiments demonstrate that both fine-tuning and knowledge distillation yield strong performance, with cross-distillation proving particularly effective in improving in-domain performance on Xeno-canto data. However, when generalizing to soundscapes, shallow fine-tuning exhibits superior performance compared to knowledge distillation, highlighting its robustness and constrained nature. Our study further investigates how to use multi-species labels, in cases where these are present but incomplete. We advocate for more comprehensive labeling practices within the animal sound community, including annotating background species and providing temporal details, to enhance the training of robust bird sound classifiers. These findings provide insights into the optimal reuse of pretrained models for advancing automatic bioacoustic recognition.

📄 PDF Abstract BibTeX arXiv:2409.15383

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationSound ClassificationTransfer Learning

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Automatic recognition of element classes and boundaries in the birdsong with variable sequences

2016-01-23 · Takuya Koumura, Kazuo Okanoya

Researches on sequential vocalization often require analysis of vocalizations in long continuous sounds. In such studies as developmental ones or studies across generations in which days or months of vocalizations must b…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Boundary DetectionGeneral Classification+2

Performance Analysis of Hybrid Quantum-Classical Convolutional Neural Networks for Audio Classification

2024-11-04 · 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT) 2024 11 · Yash Thakar, Bhuvi Ghosh, Vishma Adeshra, Kriti Srivastava

Audio signals being high-dimensional and complex pose challenges for classical machine learning techniques in terms of computation and generalization on real-world data. This paper evaluates the use of hybrid quantum-cla…

Audio ClassificationQuantum Machine Learning

Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis

2025-11-15 · Houtan Ghaffari, Lukas Rauch, Paul Devos arxiv

Research in bioacoustics, neuroscience, and linguistics often uses birdsong as a proxy to acquire knowledge across diverse areas. This requires audio models to annotate and parse the birdsong. Developing such models requ…

Self-Supervised LearningOnline ClusteringData Augmentation

Multi-Label Classifier Chains for Bird Sound

2013-04-22 · Forrest Briggs, Xiaoli Z. Fern, Jed Irvine

Bird sound data collected with unattended microphones for automatic surveys, or mobile devices for citizen science, typically contain multiple simultaneously vocalizing birds of different species. However, few works have…

General ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMulti-Label Learning

Machine Learning-based Classification of Birds through Birdsong

2022-12-09 · Yueying Chang, Richard O. Sinnott

Audio sound recognition and classification is used for many tasks and applications including human voice recognition, music recognition and audio tagging. In this paper we apply Mel Frequency Cepstral Coefficients (MFCC)…

Audio TaggingClassification