Deep High-Resolution Representation Learning for Human Pose Estimation
This is an official pytorch implementation of Deep High-Resolution Representation Learning for Human Pose Estimation. In this work, we are interested in the human pose estimation problem with a focus on learning reliable high-resolution representations. Most existing methods recover high-resolution representations from low-resolution representations produced by a high-to-low resolution network. Instead, our proposed network maintains high-resolution representations through the whole process. We start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one to form more stages, and connect the mutli-resolution subnetworks in parallel. We conduct repeated multi-scale fusions such that each of the high-to-low resolution representations receives information from other parallel representations over and over, leading to rich high-resolution representations. As a result, the predicted keypoint heatmap is potentially more accurate and spatially more precise. We empirically demonstrate the effectiveness of our network through the superior pose estimation results over two benchmark datasets: the COCO keypoint detection dataset and the MPII Human Pose dataset. The code and models have been publicly available at \url{https://github.com/leoxiaobin/deep-high-resolution-net.pytorch}.
Code (39)
Tasks
2D Human Pose Estimation2D Pose Estimation3D Human Pose Estimation3D Pose EstimationInstance SegmentationKeypoint DetectionMulti-Person Pose EstimationObject DetectionPose EstimationPose TrackingRepresentation LearningVocal Bursts Intensity PredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation
High-resolution representation is essential for achieving good performance in human pose estimation models. To obtain such features, existing works utilize high-resolution input images or fine-grained image tokens. Howev…
Pose EstimationDeep High-Resolution Representation Learning for Visual Recognition
High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-the-art frameworks first encode the inpu…
Dichotomous Image SegmentationFace AlignmentInstance SegmentationObject Detection+5HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation
Bottom-up human pose estimation methods have difficulties in predicting the correct pose for small persons due to challenges in scale variation. In this paper, we present HigherHRNet: a novel bottom-up human pose estimat…
2D Human Pose EstimationMulti-Person Pose EstimationPose EstimationPose Prediction+1Multi-Stage HRNet: Multiple Stage High-Resolution Network for Human Pose Estimation
Human pose estimation are of importance for visual understanding tasks such as action recognition and human-computer interaction. In this work, we present a Multiple Stage High-Resolution Network (Multi-Stage HRNet) to t…
Action RecognitionMulti-Person Pose EstimationPose EstimationPositionDual Super-Resolution Learning for Semantic Segmentation
Current state-of-the-art semantic segmentation methods often apply high-resolution input to attain high performance, which brings large computation budgets and limits their applications on resource-constrained devices. I…
Crack SegmentationImage Super-ResolutionPose EstimationSegmentation+2