Unsupervised 3D Human Pose Representation with Viewpoint and Pose Disentanglement
Learning a good 3D human pose representation is important for human pose related tasks, e.g. human 3D pose estimation and action recognition. Within all these problems, preserving the intrinsic pose information and adapting to view variations are two critical issues. In this work, we propose a novel Siamese denoising autoencoder to learn a 3D pose representation by disentangling the pose-dependent and view-dependent feature from the human skeleton data, in a fully unsupervised manner. These two disentangled features are utilized together as the representation of the 3D pose. To consider both the kinematic and geometric dependencies, a sequential bidirectional recursive network (SeBiReNet) is further proposed to model the human skeleton data. Extensive experiments demonstrate that the learned representation 1) preserves the intrinsic information of human pose, 2) shows good transferability across datasets and tasks. Notably, our approach achieves state-of-the-art performance on two inherently different tasks: pose denoising and unsupervised action recognition. Code and models are available at: \url{https://github.com/NIEQiang001/unsupervised-human-pose.git}
Code (1)
Tasks
3D Pose EstimationAction RecognitionDenoisingDisentanglementPose EstimationSelf-supervised Skeleton-based Action RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified Viewpoints
Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoint…
DiversityUnsupervised View-Invariant Human Posture Representation
Most recent view-invariant action recognition and performance assessment approaches rely on a large amount of annotated 3D skeleton data to extract view-invariant features. However, acquiring 3D skeleton data can be cumb…
3D Action Recognition3D Pose EstimationAction AnalysisAction Assessment+3Unsupervised Object-Centric Learning from Multiple Unspecified Viewpoints
Visual scenes are extremely diverse, not only because there are infinite possible combinations of objects and backgrounds but also because the observations of the same scene may vary greatly with the change of viewpoints…
ObjectUnsupervised Human Action Recognition with Skeletal Graph Laplacian and Self-Supervised Viewpoints Invariance
This paper presents a novel end-to-end method for the problem of skeleton-based unsupervised human action recognition. We propose a new architecture with a convolutional autoencoder that uses graph Laplacian regularizati…
Action RecognitionSkeleton Based Action RecognitionTemporal Action LocalizationUnsupervised Skeleton Based Action RecognitionTime-Aware and View-Aware Video Rendering for Unsupervised Representation Learning
The recent success in deep learning has lead to various effective representation learning methods for videos. However, the current approaches for video representation require large amount of human labeled datasets for ef…
Representation Learning