Quantification of Occlusion Handling Capability of a 3D Human Pose Estimation Framework
3D human pose estimation using monocular images is an important yet challenging task. Existing 3D pose detection methods exhibit excellent performance under normal conditions however their performance may degrade due to occlusion. Recently some occlusion aware methods have also been proposed, however, the occlusion handling capability of these networks has not yet been thoroughly investigated. In the current work, we propose an occlusion-guided 3D human pose estimation framework and quantify its occlusion handling capability by using different protocols. The proposed method estimates more accurate 3D human poses using 2D skeletons with missing joints as input. Missing joints are handled by introducing occlusion guidance that provides extra information about the absence or presence of a joint. Temporal information has also been exploited to better estimate the missing joints. A large number of experiments are performed for the quantification of occlusion handling capability of the proposed method on three publicly available datasets in various settings including random missing joints, fixed body parts missing, and complete frames missing, using mean per joint position error criterion. In addition to that, the quality of the predicted 3D poses is also evaluated using action classification performance as a criterion. 3D poses estimated by the proposed method achieved significantly improved action recognition performance in the presence of missing joints. Our experiments demonstrate the effectiveness of the proposed framework for handling the missing joints as well as quantification of the occlusion handling capability of the deep neural networks.
Code (1)
Tasks
3D Human Pose EstimationAction ClassificationAction RecognitionOcclusion HandlingPose EstimationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
GPT-4 for Occlusion Order Recovery
Occlusion remains a significant challenge for current vision models to robustly interpret complex and dense real-world images and scenes. To address this limitation and to enable accurate prediction of the occlusion orde…
3D Human Pose Estimation using Spatio-Temporal Networks with Explicit Occlusion Training
Estimating 3D poses from a monocular video is still a challenging task, despite the significant progress that has been made in recent years. Generally, the performance of existing methods drops when the target person is …
3D Human Pose EstimationMonocular 3D Human Pose EstimationPose EstimationvalidSTRIDE: Single-video based Temporally Continuous Occlusion-Robust 3D Pose Estimation
The capability to accurately estimate 3D human poses is crucial for diverse fields such as action recognition, gait recognition, and virtual/augmented reality. However, a persistent and significant challenge within this …
3D Human Pose Estimation3D Pose EstimationAction RecognitionGait Recognition+2Occlusion Handling in Generic Object Detection: A Review
The significant power of deep learning networks has led to enormous development in object detection. Over the last few years, object detector frameworks have achieved tremendous success in both accuracy and efficiency. H…
Objectobject-detectionObject DetectionOcclusion HandlingMulti-Scale Networks for 3D Human Pose Estimation with Inference Stage Optimization
Estimating 3D human poses from a monocular video is still a challenging task. Many existing methods' performance drops when the target person is occluded by other objects, or the motion is too fast/slow relative to the s…
2D Pose Estimation3D Human Pose EstimationPose EstimationPose Prediction+1