Monocular Total Capture: Posing Face, Body, and Hands in the Wild
We present the first method to capture the 3D total motion of a target person from a monocular view input. Given an image or a monocular video, our method reconstructs the motion from body, face, and fingers represented by a 3D deformable mesh model. We use an efficient representation called 3D Part Orientation Fields (POFs), to encode the 3D orientations of all body parts in the common 2D image space. POFs are predicted by a Fully Convolutional Network (FCN), along with the joint confidence maps. To train our network, we collect a new 3D human motion dataset capturing diverse total body motion of 40 subjects in a multiview system. We leverage a 3D deformable human model to reconstruct total body pose from the CNN outputs by exploiting the pose and shape prior in the model. We also present a texture-based tracking method to obtain temporally coherent motion capture output. We perform thorough quantitative evaluations including comparison with the existing body-specific and hand-specific methods, and performance analysis on camera viewpoint and human pose changes. Finally, we demonstrate the results of our total body motion capture on various challenging in-the-wild videos. Our code and newly collected human motion dataset will be publicly shared.
Code (1)
Tasks
3D Human Pose EstimationHand Pose EstimationMonocular 3D Human Pose EstimationSimilar Papers 제목 키워드 기반
Imposing Temporal Consistency on Deep Monocular Body Shape and Pose Estimation
Accurate and temporally consistent modeling of human bodies is essential for a wide range of applications, including character animation, understanding human social behavior and AR/VR interfaces. Capturing human motion a…
Pose EstimationFrankMocap: A Monocular 3D Whole-Body Pose Estimation System via Regression and Integration
Most existing monocular 3D pose estimation approaches only focus on a single body part, neglecting the fact that the essential nuance of human motion is conveyed through a concert of subtle movements of face, hands, and …
3D Human Pose Estimation3D Human Reconstruction3D Pose EstimationPose Estimation+1Monocular Real-time Full Body Capture with Inter-part Correlations
We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network archit…
3D Hand Pose EstimationComputational EfficiencyFace ModelFrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration
Although the essential nuance of human motion is often conveyed as a combination of body movements and hand gestures, the existing monocular motion capture approaches mostly focus on either body motion capture only ignor…
3D Hand Pose Estimation3D Human Reconstruction3D Pose EstimationregressionA Novel Method to Improve Quality Surface Coverage in Multi-View Capture
The depth of field of a camera is a limiting factor for applications that require taking images at a short subject-to-camera distance or using a large focal length, such as total body photography, archaeology, and other …