Hierarchical Contrastive Motion Learning for Video Action Recognition
One central question for video action recognition is how to model motion. In this paper, we present hierarchical contrastive motion learning, a new self-supervised learning framework to extract effective motion representations from raw video frames. Our approach progressively learns a hierarchy of motion features that correspond to different abstraction levels in a network. This hierarchical design bridges the semantic gap between low-level motion cues and high-level recognition tasks, and promotes the fusion of appearance and motion information at multiple levels. At each level, an explicit motion self-supervision is provided via contrastive learning to enforce the motion features at the current level to predict the future ones at the previous level. Thus, the motion features at higher levels are trained to gradually capture semantic dynamics and evolve more discriminative for action recognition. Our motion learning module is lightweight and flexible to be embedded into various backbone networks. Extensive experiments on four benchmarks show that the proposed approach consistently achieves superior results.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionContrastive LearningSelf-Supervised LearningTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
Video recognition remains an open challenge, requiring the identification of diverse content categories within videos. Mainstream approaches often perform flat classification, overlooking the intrinsic hierarchical struc…
Action RecognitionVideo RecognitionVideo UnderstandingMaCLR: Motion-aware Contrastive Learning of Representations for Videos
We present MaCLR, a novel method to explicitly perform cross-modal self-supervised video representations learning from visual and motion modalities. Compared to previous video representation learning methods that mostly …
Action DetectionAction RecognitionContrastive LearningRepresentation LearningMoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment
We aim to learn a joint representation between inertial measurement unit (IMU) signals and 2D pose sequences extracted from video, enabling accurate cross-modal retrieval, temporal synchronization, subject and body-part …
Cross-Modal RetrievalContrastive LearningAction RecognitionHierarchically Self-Supervised Transformer for Human Skeleton Representation Learning
Despite the success of fully-supervised human skeleton sequence modeling, utilizing self-supervised pre-training for skeleton sequence representation learning has been an active field because acquiring task-specific skel…
Action DetectionAction RecognitionContrastive Learningmotion prediction+1Multiscale Multimodal Transformer for Multimodal Action Recognition
While action recognition has been an active research area for several years, most existing approaches merely leverage the video modality as opposed to humans that efficiently process video and audio cues simultaneously. …
Action RecognitionAudio ClassificationMulti-modal ClassificationRepresentation Learning