Identity-Aware Hand Mesh Estimation and Personalization from RGB Images
Reconstructing 3D hand meshes from monocular RGB images has attracted increasing amount of attention due to its enormous potential applications in the field of AR/VR. Most state-of-the-art methods attempt to tackle this task in an anonymous manner. Specifically, the identity of the subject is ignored even though it is practically available in real applications where the user is unchanged in a continuous recording session. In this paper, we propose an identity-aware hand mesh estimation model, which can incorporate the identity information represented by the intrinsic shape parameters of the subject. We demonstrate the importance of the identity information by comparing the proposed identity-aware model to a baseline which treats subject anonymously. Furthermore, to handle the use case where the test subject is unseen, we propose a novel personalization pipeline to calibrate the intrinsic shape parameters using only a few unlabeled RGB images of the subject. Experiments on two large scale public datasets validate the state-of-the-art performance of our proposed method.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
AvatarMix: Identity-Preserving Cross-Avatar Composition for Outfit Personalization
Existing 3D avatar outfit transfer methods face distinct challenges: approaches that lift 2D edits to 3D often suffer from outfit or identity quality degradation, while those that separately model body and clothing layer…
HandMIM: Pose-Aware Self-Supervised Learning for 3D Hand Mesh Estimation
With an enormous number of hand images generated over time, unleashing pose knowledge from unlabeled images for supervised hand mesh estimation is an emerging yet challenging topic. To alleviate this issue, semi-supervis…
Pose EstimationregressionRepresentation LearningSelf-Supervised LearningDreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
While 2D diffusion models have achieved remarkable success in identity-preserving personalization, extending this capability to 3D assets remains a significant challenge due to the complexities of multi-view consistency …
Meta-LoRA: Meta-Learning LoRA Components for Domain-Aware ID Personalization
Recent advancements in text-to-image generative models, particularly latent diffusion models (LDMs), have demonstrated remarkable capabilities in synthesizing high-quality images from textual prompts. However, achieving …
Computational EfficiencyMeta-LearningMotion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor…
Motion SynthesisVideo Generation