Explicit Spatiotemporal Joint Relation Learning for Tracking Human Pose
We present a method for human pose tracking that is based on learning spatiotemporal relationships among joints. Beyond generating the heatmap of a joint in a given frame, our system also learns to predict the offset of the joint from a neighboring joint in the frame. Additionally, it is trained to predict the displacement of the joint from its position in the previous frame, in a manner that can account for possibly changing joint appearance, unlike optical flow. These relational cues in the spatial domain and temporal domain are inferred in a robust manner by attending only to relevant areas in the video frames. By explicitly learning and exploiting these joint relationships, our system achieves state-of-the-art performance on standard benchmarks for various pose tracking tasks including 3D body pose tracking in RGB video, 3D hand pose tracking in depth sequences, and 3D hand gesture tracking in RGB video.
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Flow EstimationPose EstimationPose TrackingPositionRelationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Constrained Dynamic Correlations in Spatiotemporal Graphs for Motion Prediction
Human motion prediction is challenging due to the complex spatiotemporal feature modeling. Among all methods, graph convolution networks (GCNs) are extensively utilized because of their superiority in explicit connection…
Human motion predictionmotion predictionRealistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling
To bridge the physical and virtual worlds for rapidly developed VR/AR applications, the ability to realistically drive 3D full-body avatars is of great significance. Although real-time body tracking with only the head-mo…
ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation
We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicit…
Human-Object Interaction DetectionHuman-Object Interaction GenerationFlow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motion: static multi-view reconstruction systems assume a rigid world, wherea…
Camera Pose EstimationScene UnderstandingGRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object Tracking
Recently, 3D multi-object tracking (MOT) has widely adopted the standard tracking-by-detection paradigm, which solves the association problem between detections and tracks. Many tracking-by-detection approaches estab…
3D Multi-Object TrackingMulti-Object TrackingObject TrackingRelation