Pose from Action: Unsupervised Learning of Pose Features based on Motion
Human actions are comprised of a sequence of poses. This makes videos of humans a rich and dense source of human poses. We propose an unsupervised method to learn pose features from videos that exploits a signal which is complementary to appearance and can be used as supervision: motion. The key idea is that humans go through poses in a predictable manner while performing actions. Hence, given two poses, it should be possible to model the motion that caused the change between them. We represent each of the poses as a feature in a CNN (Appearance ConvNet) and generate a motion encoding from optical flow maps using a separate CNN (Motion ConvNet). The data for this task is automatically generated allowing us to train without human supervision. We demonstrate the strength of the learned representation by finetuning the trained model for Pose Estimation on the FLIC dataset, for static image action recognition on PASCAL and for action recognition in videos on UCF101 and HMDB51.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction Recognition In VideosOptical Flow EstimationPose EstimationTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Estimating Motion Codes from Demonstration Videos
A motion taxonomy can encode manipulations as a binary-encoded representation, which we refer to as motion codes. These motion codes innately represent a manipulation action in an embedded space that describes the motion…
Unsupervised Learning of View-invariant Action Representations
The recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, la…
Action RecognitionRepresentation LearningTemporal Action LocalizationUnsupervised Representations Improve Supervised Learning in Speech Emotion Recognition
Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction by enabling a deeper understanding of emotional states across a wide range of applications, contributing to more empathetic an…
Emotion ClassificationEmotion RecognitionSpeech Emotion RecognitionTransfer LearningUnsupervised shape and motion analysis of 3822 cardiac 4D MRIs of UK Biobank
We perform unsupervised analysis of image-derived shape and motion features extracted from 3822 cardiac 4D MRIs of the UK Biobank. First, with a feature extraction method previously published based on deep learning model…
ClusteringDimensionality Reductionfeature selectionGeneral ClassificationTubelets: Unsupervised action proposals from spatiotemporal super-voxels
This paper considers the problem of localizing actions in videos as a sequences of bounding boxes. The objective is to generate action proposals that are likely to include the action of interest, ideally achieving high r…
Action Localization