DivaTrack: Diverse Bodies and Motions from Acceleration-Enhanced Three-Point Trackers
Full-body avatar presence is crucial for immersive social and environmental interactions in digital reality. However, current devices only provide three six degrees of freedom (DOF) poses from the headset and two controllers (i.e. three-point trackers). Because it is a highly under-constrained problem, inferring full-body pose from these inputs is challenging, especially when supporting the full range of body proportions and use cases represented by the general population. In this paper, we propose a deep learning framework, DivaTrack, which outperforms existing methods when applied to diverse body sizes and activities. We augment the sparse three-point inputs with linear accelerations from Inertial Measurement Units (IMU) to improve foot contact prediction. We then condition the otherwise ambiguous lower-body pose with the predictions of foot contact and upper-body pose in a two-stage model. We further stabilize the inferred full-body pose in a wide range of configurations by learning to blend predictions that are computed in two reference frames, each of which is designed for different types of motions. We demonstrate the effectiveness of our design on a large dataset that captures 22 subjects performing challenging locomotion for three-point tracking, including lunges, hula-hooping, and sitting. As shown in a live demo using the Meta VR headset and Xsens IMUs, our method runs in real-time while accurately tracking a user's motion when they perform a diverse set of movements.
Code (1)
Tasks
Point TrackingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Non-linear dynamics of multibody systems: a system-based approach
This paper presents causal block-diagram models to represent the equations of motion of multi-body systems in a very compact and simple closed form. Both the forward dynamics (from the forces and torques imposed at the v…
Modulating Language Models with Emotions
Generating context-aware language that embodies diverse emotions is an important step towards building empathetic NLP systems. In this paper, we propose a formulation of modulated layer normalization -- a technique inspi…
DiversityResponse GenerationIQ-VFI: Implicit Quadratic Motion Estimation for Video Frame Interpolation
Advanced video frame interpolation (VFI) algorithms approximate intermediate motions between two input frames to synthesize intermediate frame. However they struggle to handle complex scenarios with curvilinear motio…
Knowledge DistillationMotion EstimationVideo Frame InterpolationJerk-Aware Video Acceleration Magnification
Video magnification reveals subtle changes invisible to the naked eye, but such tiny yet meaningful changes are often hidden under large motions: small deformation of the muscles in doing sports, or tiny vibrations of st…
Time SeriesTime Series AnalysisFrom Universal Humanoid Control to Automatic Physically Valid Character Creation
Automatically designing virtual humans and humanoids holds great potential in aiding the character creation process in games, movies, and robots. In some cases, a character creator may wish to design a humanoid body cust…
Humanoid Controlvalid