Towards High Performance Human Keypoint Detection
Human keypoint detection from a single image is very challenging due to occlusion, blur, illumination and scale variance. In this paper, we address this problem from three aspects by devising an efficient network structure, proposing three effective training strategies, and exploiting four useful postprocessing techniques. First, we find that context information plays an important role in reasoning human body configuration and invisible keypoints. Inspired by this, we propose a cascaded context mixer (CCM), which efficiently integrates spatial and channel context information and progressively refines them. Then, to maximize CCM's representation capability, we develop a hard-negative person detection mining strategy and a joint-training strategy by exploiting abundant unlabeled data. It enables CCM to learn discriminative features from massive diverse poses. Third, we present several sub-pixel refinement techniques for postprocessing keypoint predictions to improve detection accuracy. Extensive experiments on the MS COCO keypoint detection benchmark demonstrate the superiority of the proposed method over representative state-of-the-art (SOTA) methods. Our single model achieves comparable performance with the winner of the 2018 COCO Keypoint Detection Challenge. The final ensemble model sets a new SOTA on this benchmark.
Code (1)
Tasks
Human DetectionKeypoint DetectionPose EstimationVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
Human keypoint detection for close proximity human-robot interaction
We study the performance of state-of-the-art human keypoint detectors in the context of close proximity human-robot interaction. The detection in this scenario is specific in that only a subset of body parts such as hand…
Keypoint DetectionImproving Multi-Person Pose Tracking with A Confidence Network
Human pose estimation and tracking are fundamental tasks for understanding human behaviors in videos. Existing top-down framework-based methods usually perform three-stage tasks: human detection, pose estimation and trac…
Human DetectionPose EstimationPose TrackingInterspecies Knowledge Transfer for Facial Keypoint Detection
We present a method for localizing facial keypoints on animals by transferring knowledge gained from human faces. Instead of directly finetuning a network trained to detect keypoints on human faces to animal faces (which…
Human DetectionKeypoint DetectionTransfer LearningExplicit Box Detection Unifies End-to-End Multi-Person Pose Estimation
This paper presents a novel end-to-end framework with Explicit box Detection for multi-person Pose estimation, called ED-Pose, where it unifies the contextual learning between human-level (global) and keypoint-level (loc…
2D Human Pose EstimationDecoderHuman DetectionKeypoint Detection+3HOKEM: Human and Object Keypoint-based Extension Module for Human-Object Interaction Detection
Human-object interaction (HOI) detection for capturing relationships between humans and objects is an important task in the semantic understanding of images. When processing human and object keypoints extracted from an i…
Human-Object Interaction DetectionObject