Audio Self-supervised Learning: A Survey
Inspired by the humans' cognitive ability to generalise knowledge and skills, Self-Supervised Learning (SSL) targets at discovering general representations from large-scale data without requiring human annotations, which is an expensive and time consuming task. Its success in the fields of computer vision and natural language processing have prompted its recent adoption into the field of audio and speech processing. Comprehensive reviews summarising the knowledge in audio SSL are currently missing. To fill this gap, in the present work, we provide an overview of the SSL methods used for audio and speech processing applications. Herein, we also summarise the empirical works that exploit the audio modality in multi-modal SSL frameworks, and the existing suitable benchmarks to evaluate the power of SSL in the computer audition domain. Finally, we discuss some open problems and point out the future directions on the development of audio SSL.
Code (0)
등록된 구현이 없습니다.
Tasks
Self-Supervised LearningSurveySimilar Papers 제목 키워드 기반
TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch
TorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing well-designed, easy-to-use, and perform…
Self-Supervised LearningSpeech Enhancementspeech-recognitionSpeech RecognitionFrom Waveforms to Pixels: A Survey on Audio-Visual Segmentation
Audio-Visual Segmentation (AVS) aims to identify and segment sound-producing objects in videos by leveraging both visual and audio modalities. It has emerged as a significant research area in multimodal perception, enabl…
Few-Shot LearningAudio Barlow Twins: Self-Supervised Audio Representation Learning
The Barlow Twins self-supervised learning objective requires neither negative samples or asymmetric learning updates, achieving results on a par with the current state-of-the-art within Computer Vision. As such, we prese…
Environmental Sound ClassificationEvent DetectionRepresentation LearningSelf-Supervised LearningLearning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision
The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and a…
Acoustic Scene ClassificationAction RecognitionScene ClassificationSelf-Supervised Learning+1Conformer-Based Self-Supervised Learning for Non-Speech Audio Tasks
Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few…
Audio ClassificationRepresentation LearningSelf-Supervised LearningSpeech Representation Learning