Improving label efficiency through multi-task learning on auditory data
Collecting high-quality, large scale datasets typically requires significant resources. The aim of the present work is to improve the label efficiency of large neural networks operating on audio data through multitask learning with self-supervised tasks on unlabeled data. To this end, we trained an end-to-end audio feature extractor based on WaveNet that feeds into simple, yet versatile task-specific neural networks. We describe three self-supervised learning tasks that can operate on any large, unlabeled audio corpus. We demonstrate that, in a scenario with limited labeled training data, one can significantly improve the performance of a supervised classification task by simultaneously training it with these additional self-supervised tasks. We show that one can improve performance on a diverse sound events classification task by nearly 6\% when jointly trained with up to three distinct self-supervised tasks. This improvement scales with the number of additional auxiliary tasks as well as the amount of unsupervised data. We also show that incorporating data augmentation into our multitask setting leads to even further gains in performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationMulti-Task LearningSelf-Supervised LearningSimilar Papers 제목 키워드 기반
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
Large language models (LLMs) have shown exceptional versatility in natural language processing, prompting recent efforts to extend their multimodal capabilities to speech processing through the development of audio large…
speech-recognitionSpeech RecognitionJointly Learning Visual and Auditory Speech Representations from Raw Data
We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…
Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often l…
Spiking-LEAF: A Learnable Auditory front-end for Spiking Neural Networks
Brain-inspired spiking neural networks (SNNs) have demonstrated great potential for temporal signal processing. However, their performance in speech processing remains limited due to the lack of an effective auditory fro…
Keyword SpottingSpeaker IdentificationHow to train your ears: Auditory-model emulation for large-dynamic-range inputs and mild-to-severe hearing losses
Advanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and …
Speech Enhancement