paper-with-me

Papers

HSTFormer: Hierarchical Spatial-Temporal Transformers for 3D Human Pose Estimation

2023-01-18 · Xiaoye Qian, YouBao Tang, Ning Zhang, Mei Han, Jing Xiao, Ming-Chun Huang, Ruei-Sung Lin

Transformer-based approaches have been successfully proposed for 3D human pose estimation (HPE) from 2D pose sequence and achieved state-of-the-art (SOTA) performance. However, current SOTAs have difficulties in modeling spatial-temporal correlations of joints at different levels simultaneously. This is due to the poses' spatial-temporal complexity. Poses move at various speeds temporarily with various joints and body-parts movement spatially. Hence, a cookie-cutter transformer is non-adaptable and can hardly meet the "in-the-wild" requirement. To mitigate this issue, we propose Hierarchical Spatial-Temporal transFormers (HSTFormer) to capture multi-level joints' spatial-temporal correlations from local to global gradually for accurate 3D HPE. HSTFormer consists of four transformer encoders (TEs) and a fusion module. To the best of our knowledge, HSTFormer is the first to study hierarchical TEs with multi-level fusion. Extensive experiments on three datasets (i.e., Human3.6M, MPI-INF-3DHP, and HumanEva) demonstrate that HSTFormer achieves competitive and consistent performance on benchmarks with various scales and difficulties. Specifically, it surpasses recent SOTAs on the challenging MPI-INF-3DHP dataset and small-scale HumanEva dataset, with a highly generalized systematic approach. The code is available at: https://github.com/qianxiaoye825/HSTFormer.

📄 PDF Abstract BibTeX arXiv:2301.07322

Code (0)

등록된 구현이 없습니다.

Tasks

3D Human Pose EstimationPose Estimation

Similar Papers 제목 키워드 기반

Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser

2024-03-07 · Qingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao 외

Recently, diffusion-based methods for monocular 3D human pose estimation have achieved state-of-the-art (SOTA) performance by directly regressing the 3D joint coordinates from the 2D pose sequence. Although some methods …

3D Human Pose EstimationDisentanglementMonocular 3D Human Pose EstimationMulti-Hypotheses 3D Human Pose Estimation+1

Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

2024-05-27 · Ruizhi Shao, Youxin Pang, Zerong Zheng, Jingxiang Sun 외

We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correla…

Video Generation

Hierarchical Separable Video Transformer for Snapshot Compressive Imaging

2024-07-16 · Ping Wang, Yulun Zhang, Lishun Wang, Xin Yuan

Transformers have achieved the state-of-the-art performance on solving the inverse problem of Snapshot Compressive Imaging (SCI) for video, whose ill-posedness is rooted in the mixed degradation of spatial masking and te…

Inductive BiasLong-range modeling

Spatial-Temporal Interplay in Human Mobility: A Hierarchical Reinforcement Learning Approach with Hypergraph Representation

2023-12-25 · Zhaofan Zhang, Yanan Xiao, Lu Jiang, Dingqi Yang 외

In the realm of human mobility, the decision-making process for selecting the next-visit location is intricately influenced by a trade-off between spatial and temporal constraints, which are reflective of individual need…

Decision MakingHierarchical Reinforcement Learninghypergraph embedding

Point Primitive Transformer for Long-Term 4D Point Cloud Video Understanding

2022-07-30 · Hao Wen, Yunze Liu, Jingwei Huang, Bo Duan 외

This paper proposes a 4D backbone for long-term point cloud video understanding. A typical way to capture spatial-temporal context is using 4Dconv or transformer without hierarchy. However, those methods are neither effe…

point cloud video understandingVideo Understanding