BYOL-S: Learning Self-supervised Speech Representations by Bootstrapping
Methods for extracting audio and speech features have been studied since pioneering work on spectrum analysis decades ago. Recent efforts are guided by the ambition to develop general-purpose audio representations. For example, deep neural networks can extract optimal embeddings if they are trained on large audio datasets. This work extends existing methods based on self-supervised learning by bootstrapping, proposes various encoder architectures, and explores the effects of using different pre-training datasets. Lastly, we present a novel training framework to come up with a hybrid audio representation, which combines handcrafted and data-driven learned audio features. All the proposed representations were evaluated within the HEAR NeurIPS 2021 challenge for auditory scene classification and timestamp detection tasks. Our results indicate that the hybrid model with a convolutional transformer as the encoder yields superior performance in most HEAR challenge tasks.
Code (1)
Tasks
Scene ClassificationSelf-Supervised LearningSimilar Papers 제목 키워드 기반
Cross-view Self-Supervised Learning on Heterogeneous Graph Neural Network via Bootstrapping
Heterogeneous graph neural networks can represent information of heterogeneous graphs with excellent ability. Recently, self-supervised learning manner is researched which learns the unique expression of a graph through …
Contrastive LearningGraph Neural NetworkSelf-Supervised LearningCompressive Visual Representations
Learning effective visual representations that generalize well without human supervision is a fundamental problem in order to apply Machine Learning to a wide variety of tasks. Recently, two families of self-supervised m…
Contrastive LearningImage ClassificationLinear evaluationSelf-Supervised Image ClassificationA Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning
Learning a good representation is a crucial challenge for Reinforcement Learning (RL) agents. Self-predictive learning provides means to jointly learn a latent representation and dynamics model by bootstrapping from futu…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Representation LearningRun Away From your Teacher: Understanding BYOL by a Novel Self-Supervised Approach
Recently, a newly proposed self-supervised framework Bootstrap Your Own Latent (BYOL) seriously challenges the necessity of negative samples in contrastive learning frameworks. BYOL works like a charm despite the fact th…
Contrastive LearningSelf-Supervised LearningRun Away From your Teacher: a New Self-Supervised Approach Solving the Puzzle of BYOL
Recently, a newly proposed self-supervised framework Bootstrap Your Own Latent (BYOL) seriously challenges the necessity of negative samples in contrastive-based learning frameworks. BYOL works like a charm despite the f…
Self-Supervised Learning