Inferring Implicit 3D Representations from Human Figures on Pictorial Maps
In this work, we present an automated workflow to bring human figures, one of the most frequently appearing entities on pictorial maps, to the third dimension. Our workflow is based on training data and neural networks for single-view 3D reconstruction of real humans from photos. We first let a network consisting of fully connected layers estimate the depth coordinate of 2D pose points. The gained 3D pose points are inputted together with 2D masks of body parts into a deep implicit surface network to infer 3D signed distance fields (SDFs). By assembling all body parts, we derive 2D depth images and body part masks of the whole figure for different views, which are fed into a fully convolutional network to predict UV images. These UV images and the texture for the given perspective are inserted into a generative network to inpaint the textures for the other views. The textures are enhanced by a cartoonization network and facial details are resynthesized by an autoencoder. Finally, the generated textures are assigned to the inferred body parts in a ray marcher. We test our workflow with 12 pictorial human figures after having validated several network configurations. The created 3D models look generally promising, especially when considering the challenges of silhouette-based 3D recovery and real-time rendering of the implicit SDFs. Further improvement is needed to reduce gaps between the body parts and to add pictorial details to the textures. Overall, the constructed figures may be used for animation and storytelling in digital 3D maps.
Code (0)
등록된 구현이 없습니다.
Tasks
3D ReconstructionSingle-View 3D ReconstructionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Annotating shadows, highlights and faces: the contribution of a 'human in the loop' for digital art history
While automatic computational techniques appear to reveal novel insights in digital art history, a complementary approach seems to get less attention: that of human annotation. We argue and exemplify that a 'human in the…
Artificial Phantasia: Emergent Mental Imagery in Large Language Models
Can visual imagery be driven solely by language? This idea goes against cognitive science's traditional view that visual mental imagery is only possible through pictorial representations. Large Language Models (LLMs) pro…
Tinkering Under the Hood: Interactive Zero-Shot Learning with Net Surgery
We consider the task of visual net surgery, in which a CNN can be reconfigured without extra data to recognize novel concepts that may be omitted from the training set. While most prior work make use of linguistic cues f…
Novel ConceptsZero-Shot LearningFAMOUS: High-Fidelity Monocular 3D Human Digitization Using View Synthesis
The advancement in deep implicit modeling and articulated models has significantly enhanced the process of digitizing human figures in 3D from just a single image. While state-of-the-art methods have greatly improved geo…
Multiple human pose estimation with temporally consistent 3d pictorial structures
Multiple human 3D pose estimation from multiple camera views is a challenging task in unconstrained environments. Each individual has to be matched across each view and then the body pose has to be estimated. Additionall…
3D Multi-Person Pose Estimation3D Pose EstimationPose Estimation