Self-Supervised Action Recognition
6개 벤치마크 · 논문 51편 · 이 태스크의 논문 보기 →
Benchmarks
UCF101
HMDB51
HMDB51 (finetuned)
UCF101 (finetuned)
Kinetics-600
Kinetics-400
Most implemented
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Contrastive Multiview Coding
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning
Spatiotemporal Contrastive Video Representation Learning
Papers
Feature Hallucination for Self-supervised Action Recognition
Understanding human actions in videos requires more than raw pixel analysis; it relies on high-level semantic reasoning and effective integration of multimodal features. We propose a deep translational action recognition…
Action RecognitionHallucinationobject-detectionObject Detection+3Cross-Model Cross-Stream Learning for Self-Supervised Human Action Recognition
Considering the instance-level discriminative ability, contrastive learning methods, including MoCo and SimCLR, have been adapted from the original image representation learning task to solve the self-supervised skeleton…
Action RecognitionContrastive LearningEnsemble LearningPseudo Label+6A Large-Scale Analysis on Self-Supervised Video Representation Learning
Self-supervised learning is an effective way for label-free model pre-training, especially in the video domain where labeling is expensive. Existing self-supervised works in the video domain use varying experimental setu…
BenchmarkingRepresentation LearningSelf-Supervised Action RecognitionSelf-Supervised LearningPart Aware Contrastive Learning for Self-Supervised Action Recognition
In recent years, remarkable results have been achieved in self-supervised action recognition using skeleton sequences with contrastive learning. It has been observed that the semantic distinction of human action features…
Action RecognitionContrastive LearningData AugmentationRepresentation Learning+3VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
Scale is the primary factor for building a powerful foundation model that could well generalize to a variety of downstream tasks. However, it is still challenging to train video foundation models with billions of paramet…
Action ClassificationAction RecognitionAction Recognition In VideosDecoder+3Self-supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences
Self-supervised learning has demonstrated remarkable capability in representation learning for skeleton-based action recognition. Existing methods mainly focus on applying global data augmentation to generate different v…
Action RecognitionContrastive LearningData AugmentationRepresentation Learning+5