Deep representation learning for human motion prediction and classification
Generative models of 3D human motion are often restricted to a small number of activities and can therefore not generalize well to novel movements or applications. In this work we propose a deep learning framework for human motion capture data that learns a generic representation from a large corpus of motion capture data and generalizes well to new, unseen, motions. Using an encoding-decoding network that learns to predict future 3D poses from the most recent past, we extract a feature representation of human motion. Most work on deep learning for sequence prediction focuses on video and speech. Since skeletal data has a different structure, we present and evaluate different network architectures that make different assumptions about time dependencies and limb correlations. To quantify the learned features, we use the output of different layers for action classification and visualize the receptive fields of the network units. Our method outperforms the recent state of the art in skeletal motion prediction even though these use action specific training data. Our results show that deep feedforward networks, trained from a generic mocap database, can successfully be used for feature extraction from human motion data and that this representation can be used as a foundation for classification and prediction.
Code (0)
등록된 구현이 없습니다.
Tasks
Action ClassificationClassificationGeneral ClassificationHuman motion predictionmotion predictionPredictionRepresentation LearningSimilar Papers 제목 키워드 기반
Human Motion Prediction via Pattern Completion in Latent Representation Space
Inspired by ideas in cognitive science, we propose a novel and general approach to solve human motion understanding via pattern completion on a learned latent representation space. Our model outperforms current state-of-…
Action ClassificationGeneral ClassificationHuman motion predictionMotion Generation+3DynamoNet: Dynamic Action and Motion Network
In this paper, we are interested in self-supervised learning the motion cues in videos using dynamic motion filters for a better motion representation to finally boost human action recognition in particular. Thus far, th…
Action RecognitionClassificationGeneral ClassificationMulti-Task Learning+3Distributed Representations of Emotion Categories in Emotion Space
Emotion category is usually divided into different ones by human beings, but it is indeed difficult to clearly distinguish and define the boundaries between different emotion categories. The existing studies working on e…
Emotion ClassificationHuman Motion Prediction via Learning Local Structure Representations and Temporal Dependencies
Human motion prediction from motion capture data is a classical problem in the computer vision, and conventional methods take the holistic human body as input. These methods ignore the fact that, in various human activit…
Human motion predictionmotion predictionDecompose More and Aggregate Better: Two Closer Looks at Frequency Representation Learning for Human Motion Prediction
Encouraged by the effectiveness of encoding temporal dynamics within the frequency domain, recent human motion prediction systems prefer to first convert the motion representation from the original pose space into th…
Human motion predictionmotion predictionPredictionRepresentation Learning