paper-with-me

Papers

Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing

2025-01-10 · Eklavya Sarkar, Mathew Magimai. -Doss

Self-supervised learning (SSL) foundation models have emerged as powerful, domain-agnostic, general-purpose feature extractors applicable to a wide range of tasks. Such models pre-trained on human speech have demonstrated high transferability for bioacoustic processing. This paper investigates (i) whether SSL models pre-trained directly on animal vocalizations offer a significant advantage over those pre-trained on speech, and (ii) whether fine-tuning speech-pretrained models on automatic speech recognition (ASR) tasks can enhance bioacoustic classification. We conduct a comparative analysis using three diverse bioacoustic datasets and two different bioacoustic tasks. Results indicate that pre-training on bioacoustic data provides only marginal improvements over speech-pretrained models, with comparable performance in most scenarios. Fine-tuning on ASR tasks yields mixed outcomes, suggesting that the general-purpose representations learned during SSL pre-training are already well-suited for bioacoustic tasks. These findings highlight the robustness of speech-pretrained SSL models for bioacoustics and imply that extensive fine-tuning may not be necessary for optimal performance.

📄 PDF Abstract BibTeX arXiv:2501.05987

Code (1)

idiap/ssl-human-animal 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Unsupervised Representations Improve Supervised Learning in Speech Emotion Recognition

2023-09-22 · Amirali Soltani Tehrani, Niloufar Faridani, Ramin Toosi

Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction by enabling a deeper understanding of emotional states across a wide range of applications, contributing to more empathetic an…

Emotion ClassificationEmotion RecognitionSpeech Emotion RecognitionTransfer Learning

Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

2022-06-25 · Yihan Wu, Xi Wang, Shaofei Zhang, Lei He 외

Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of lab…

Contrastive LearningDeep ClusteringExpressive Speech SynthesisRepresentation Learning+1

Refining Self-Supervised Learnt Speech Representation using Brain Activations

2024-06-12 · Hengyu Li, Kangdi Mei, Zhaoci Liu, Yang Ai 외

It was shown in literature that speech representations extracted by self-supervised pre-trained models exhibit similarities with brain activations of human for speech perception and fine-tuning speech representation mode…

Automatic Speech RecognitionSpeaker Verificationspeech-recognitionSpeech Recognition

Pretrained self-supervised speech models can recognize unseen consonants

2026-06-10 · Chihiro Taguchi, Éric Le Ferrand, Hirosi Nakagawa, Hitomi Ono 외 arxiv

Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their training data are heavily skewed toward hig…

Speech Recognition

Neural Representations for Modeling Variation in Speech

2020-11-25 · Martijn Bartelds, Wietse de Vries, Faraz Sanal, Caitlin Richter 외

Variation in speech is often quantified by comparing phonetic transcriptions of the same utterance. However, manually transcribing speech is time-consuming and error prone. As an alternative, therefore, we investigate th…