An End-to-end Framework for Unconstrained Monocular 3D Hand Pose Estimation
This work addresses the challenging problem of unconstrained 3D hand pose estimation using monocular RGB images. Most of the existing approaches assume some prior knowledge of hand (such as hand locations and side information) is available for 3D hand pose estimation. This restricts their use in unconstrained environments. We, therefore, present an end-to-end framework that robustly predicts hand prior information and accurately infers 3D hand pose by learning ConvNet models while only using keypoint annotations. To achieve robustness, the proposed framework uses a novel keypoint-based method to simultaneously predict hand regions and side labels, unlike existing methods that suffer from background color confusion caused by using segmentation or detection-based technology. Moreover, inspired by the biological structure of the human hand, we introduce two geometric constraints directly into the 3D coordinates prediction that further improves its performance in a weakly-supervised training. Experimental results show that our proposed framework not only performs robustly on unconstrained setting, but also outperforms the state-of-art methods on standard benchmark datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Hand Pose EstimationHand Pose EstimationPose EstimationSimilar Papers 제목 키워드 기반
Region Deformer Networks for Unsupervised Depth Estimation from Unconstrained Monocular Videos
While learning based depth estimation from images/videos has achieved substantial progress, there still exist intrinsic limitations. Supervised methods are limited by a small amount of ground truth or labeled data and un…
Depth EstimationHandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control
We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand annotations from multi-view or marker-based…
Video GenerationUnconstrained Monocular 3D Human Pose Estimation by Action Detection and Cross-Modality Regression Forest
This work addresses the challenging problem of unconstrained 3D human pose estimation (HPE) from a novel perspective. Existing approaches struggle to operate in realistic applications, mainly due to their scene-dependent…
2D Pose Estimation3D Human Pose EstimationAction DetectionMonocular 3D Human Pose Estimation+3Towards unconstrained joint hand-object reconstruction from RGB videos
Our work aims to obtain 3D reconstruction of hands and manipulated objects from monocular videos. Reconstructing hand-object manipulations holds a great potential for robotics and learning from human demonstrations. The …
3D Hand Pose Estimation3D Reconstructionhand-object poseHand Pose Estimation+7Toward a Real-Time Framework for Accurate Monocular 3D Human Pose Estimation with Geometric Priors
Monocular 3D human pose estimation remains a challenging and ill-posed problem, particularly in real-time settings and unconstrained environments. While direct imageto-3D approaches require large annotated datasets and h…
Monocular 3D Human Pose EstimationKeypoint Detection3D Pose Estimation