Papers Self-Supervised Audio Classification
“Self-Supervised Audio Classification” 태그가 달린 논문 9편 · 필터 해제
Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed t…
Audio ClassificationAudio TaggingClassificationEnvironmental Sound Classification+9ATST: Audio Representation Learning with Teacher-Student Transformer
Self-supervised learning (SSL) learns knowledge from a large amount of unlabeled data, and then transfers the knowledge to a specific problem with a limited number of labeled data. SSL has achieved promising results in v…
Audio ClassificationInstrument RecognitionRepresentation LearningSelf-Supervised Audio Classification+3Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity
We present CrissCross, a self-supervised framework for learning audio-visual representations. A novel notion is introduced in our framework whereby in addition to learning the intra-modal and standard 'synchronous' cross…
Audio ClassificationRetrievalSelf-Supervised Action RecognitionSelf-Supervised Audio Classification+4Broaden Your Views for Self-Supervised Video Learning
Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by image techniques, where these two views …
Audio ClassificationOptical Flow EstimationRepresentation LearningSelf-Supervised Action Recognition+2Self-Supervised MultiModal Versatile Networks
Videos are a rich source of multi-modal supervision. In this work, we learn representations using self-supervision by leveraging three modalities naturally present in videos: visual, audio and language streams. To this e…
Action Recognition In VideosAudio ClassificationSelf-Supervised Action RecognitionSelf-Supervised Audio ClassificationAudio-Visual Instance Discrimination with Cross-Modal Agreement
We present a self-supervised learning approach to learn audio-visual representations from video and audio. Our method uses contrastive learning for cross-modal discrimination of video from audio and vice-versa. We show t…
Action RecognitionAudio ClassificationContrastive LearningSelf-Supervised Action Recognition+2Self-Supervised Learning by Cross-Modal Audio-Video Clustering
Visual and audio modalities are highly correlated, yet they contain different information. Their strong correlation makes it possible to predict the semantics of one from the other with good accuracy. Their intrinsic dif…
Action RecognitionAudio ClassificationClusteringDeep Clustering+4Putting An End to End-to-End: Gradient-Isolated Learning of Representations
We propose a novel deep learning method for local self-supervised representation learning that does not require labels nor end-to-end backpropagation but exploits the natural order in data instead. Inspired by the observ…
Representation LearningSelf-Supervised Audio ClassificationSelf-Supervised Image ClassificationSelf-Supervised LearningCooperative Learning of Audio and Video Models from Self-Supervised Synchronization
There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised te…
Action RecognitionAudio ClassificationSelf-Supervised Action RecognitionSelf-Supervised Audio Classification+1