Unsupervised Motion Representation Enhanced Network for Action Recognition
Learning reliable motion representation between consecutive frames, such as optical flow, has proven to have great promotion to video understanding. However, the TV-L1 method, an effective optical flow solver, is time-consuming and expensive in storage for caching the extracted optical flow. To fill the gap, we propose UF-TSN, a novel end-to-end action recognition approach enhanced with an embedded lightweight unsupervised optical flow estimator. UF-TSN estimates motion cues from adjacent frames in a coarse-to-fine manner and focuses on small displacement for each level by extracting pyramid of feature and warping one to the other according to the estimated flow of the last level. Due to the lack of labeled motion for action datasets, we constrain the flow prediction with multi-scale photometric consistency and edge-aware smoothness. Compared with state-of-the-art unsupervised motion representation learning methods, our model achieves better accuracy while maintaining efficiency, which is competitive with some supervised or more complicated approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionOptical Flow EstimationRepresentation LearningVideo UnderstandingSimilar Papers 제목 키워드 기반
Unsupervised Representations Improve Supervised Learning in Speech Emotion Recognition
Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction by enabling a deeper understanding of emotional states across a wide range of applications, contributing to more empathetic an…
Emotion ClassificationEmotion RecognitionSpeech Emotion RecognitionTransfer LearningUnsupervised Learning of View-invariant Action Representations
The recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, la…
Action RecognitionRepresentation LearningTemporal Action LocalizationKnowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stic…
counterfactualDomain AdaptationEmotion RecognitionUnsupervised Domain AdaptationActionlet-Dependent Contrastive Learning for Unsupervised Skeleton-Based Action Recognition
The self-supervised pretraining paradigm has achieved great success in skeleton-based action recognition. However, these methods treat the motion and static parts equally, and lack an adaptive design for different parts,…
Action RecognitionContrastive LearningSelf-supervised Skeleton-based Action RecognitionSkeleton Based Action Recognition+1Pose from Action: Unsupervised Learning of Pose Features based on Motion
Human actions are comprised of a sequence of poses. This makes videos of humans a rich and dense source of human poses. We propose an unsupervised method to learn pose features from videos that exploits a signal which is…
Action RecognitionAction Recognition In VideosOptical Flow EstimationPose Estimation+1