Papers Self-Supervised Action Recognition Linear
“Self-Supervised Action Recognition Linear” 태그가 달린 논문 11편 · 필터 해제
Language-based Action Concept Spaces Improve Video Self-Supervised Learning
Recent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domains with minimal supervision remains an open problem. W…
Action RecognitionConcept AlignmentSelf-Supervised Action Recognition LinearSelf-Supervised LearningLearning Video Representations from Large Language Models
We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create auto…
Action ClassificationAction RecognitionDiversityEgocentric Activity Recognition+2XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning
We present XKD, a novel self-supervised framework to learn meaningful representations from unlabelled videos. XKD is trained with two pseudo objectives. First, masked data reconstruction is performed to learn modality-sp…
Action ClassificationClassificationKnowledge DistillationRepresentation Learning+4VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets. In this paper, we show that video masked autoencoders (VideoMAE) are data-e…
4kAction ClassificationAction RecognitionSelf-Supervised Action Recognition+3Self-supervised Video Transformer
In this paper, we propose self-supervised training for video transformers using unlabeled video data. From a given video, we create local and global spatiotemporal views with varying spatial sizes and frame rates. Our se…
Action ClassificationAction RecognitionAction Recognition In VideosSelf-Supervised Action Recognition LinearVideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
MoCo is effective for unsupervised image representation learning. In this paper, we propose VideoMoCo for unsupervised video representation learning. Given a video sequence as an input sample, we improve the temporal fea…
Action RecognitionContrastive LearningRepresentation LearningSelf-Supervised Action Recognition LinearVi2CLR: Video and Image for Visual Contrastive Learning of Representation
In this paper, we introduce a novel self-supervised visual representation learning method which understands both images and videos in a joint learning fashion. The proposed neural network architecture and objectives …
Action RecognitionClusteringContrastive LearningRepresentation Learning+1Contrast and Order Representations for Video Self-Supervised Learning
This paper studies the problem of learning self-supervised representations on videos. In contrast to image modality that only requires appearance information on objects or scenes, video needs to further explore the r…
Action RecognitionSelf-Supervised Action Recognition LinearSelf-Supervised LearningSelf-supervised Co-training for Video Representation Learning
The objective of this paper is visual-only self-supervised video representation learning. We make the following contributions: (i) we investigate the benefit of adding semantic-class positives to instance-based Info Nois…
Action RecognitionContrastive LearningOptical Flow EstimationRepresentation Learning+4Spatiotemporal Contrastive Video Representation Learning
We present a self-supervised Contrastive Video Representation Learning (CVRL) method to learn spatiotemporal visual representations from unlabeled videos. Our representations are learned using a contrastive loss, where t…
Action RecognitionContrastive LearningData AugmentationRepresentation Learning+4Video Representation Learning by Dense Predictive Coding
The objective of this paper is self-supervised learning of spatio-temporal embeddings from video, suitable for human action recognition. We make three contributions: First, we introduce the Dense Predictive Coding (DPC) …
Action RecognitionRepresentation LearningSelf-Supervised Action RecognitionSelf-Supervised Action Recognition Linear+2