PanopTOP: a framework for generating viewpoint-invariant human pose estimation datasets
Human pose estimation (HPE) from RGB and depth images has recently experienced a push for viewpoint-invariant and scale-invariant pose retrieval methods. Current methods fail to generalize to unconventional viewpoints due to the lack of viewpoint-invariant data at training time. Existing datasets do not provide multiple-viewpoint observations and mostly focus on frontal views. In this work, we introduce PanopTOP, a fully automatic framework for the generation of semi-synthetic RGB and depth samples with 2D and 3D ground truth of pedestrian poses from multiple arbitrary viewpoints. Starting from the Panoptic Dataset [15], we use the PanopTOP framework to generate the PanopTOP31K dataset, consisting of 31K images from 23 different subjects recorded from diverse and challenging viewpoints, also including the top-view. Finally, we provide baseline results and cross-validation tests for our dataset, demonstrating how it is possible to generalize from the semi-synthetic to the real-world domain. The dataset and the code will be made publicly available upon acceptance.
Code (1)
Tasks
Pose EstimationPose RetrievalRetrievalSimilar Papers 제목 키워드 기반
View Invariant 3D Human Pose Estimation
The recent success of deep networks has significantly advanced 3D human pose estimation from 2D images. The diversity of capturing viewpoints and the flexibility of the human poses, however, remain some significant chall…
3D Human Pose Estimation3D Pose EstimationDiversityPose EstimationDECA: Deep viewpoint-Equivariant human pose estimation using Capsule Autoencoders
Human Pose Estimation (HPE) aims at retrieving the 3D position of human joints from images or videos. We show that current 3D HPE methods suffer a lack of viewpoint equivariance, namely they tend to fail or perform poorl…
3D Human Pose EstimationMonocular 3D Human Pose EstimationPose EstimationTowards Viewpoint Invariant 3D Human Pose Estimation
We propose a viewpoint invariant model for 3D human pose estimation from a single depth image. To achieve this, our discriminative model embeds local regions into a learned viewpoint invariant feature space. Formulated a…
3D Human Pose EstimationMulti-Task LearningPose EstimationView-invariant action recognition
Human action recognition is an important problem in computer vision. It has a wide range of applications in surveillance, human-computer interaction, augmented reality, video indexing, and retrieval. The varying pattern …
Action RecognitionRetrievalTemporal Action LocalizationMoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement
3D human pose estimation is a key enabling technology for applications such as healthcare monitoring, human-robot collaboration, and immersive gaming, but real-world deployment remains challenged by viewpoint variations.…
3D Human Pose Estimation