paper-with-me

홈 › Papers

Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition

2017-11-16 · Jian Liu, Naveed Akhtar, Ajmal Mian

Human skeleton joints are popular for action analysis since they can be easily extracted from videos to discard background noises. However, current skeleton representations do not fully benefit from machine learning with CNNs. We propose "Skepxels" a spatio-temporal representation for skeleton sequences to fully exploit the "local" correlations between joints using the 2D convolution kernels of CNN. We transform skeleton videos into images of flexible dimensions using Skepxels and develop a CNN-based framework for effective human action recognition using the resulting images. Skepxels encode rich spatio-temporal information about the skeleton joints in the frames by maximizing a unique distance metric, defined collaboratively over the distinct joint arrangements used in the skeletal image. Moreover, they are flexible in encoding compound semantic notions such as location and speed of the joints. The proposed action recognition exploits the representation in a hierarchical manner by first capturing the micro-temporal relations between the skeleton joints with the Skepxels and then exploiting their macro-temporal relations by computing the Fourier Temporal Pyramids over the CNN features of the skeletal images. We extend the Inception-ResNet CNN architecture with the proposed method and improve the state-of-the-art accuracy by 4.4% on the large scale NTU human activity dataset. On the medium-sized N-UCLA and UTH-MHAD datasets, our method outperforms the existing results by 5.7% and 9.3% respectively.

📄 PDF Abstract BibTeX arXiv:1711.05941

Code (0)

등록된 구현이 없습니다.

Tasks

Action AnalysisAction RecognitionTemporal Action Localization

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Spatio-temporal Tendency Reasoning for Human Body Pose and Shape Estimation from Videos

2022-10-07 · Boyang Zhang, Suping Wu, Hu Cao, Kehua Ma 외

In this paper, we present a spatio-temporal tendency reasoning (STR) network for recovering human body pose and shape from videos. Previous approaches have focused on how to extend 3D human datasets and temporal-based le…

3D Human Pose EstimationTemporal Sequences

Sparse Coding for Learning Interpretable Spatio-Temporal Primitives

2010-12-01 · NeurIPS 2010 12 · Taehwan Kim, Gregory Shakhnarovich, Raquel Urtasun

Sparse coding has recently become a popular approach in computer vision to learn dictionaries of natural images. In this paper we extend sparse coding to learn interpretable spatio-temporal primitives of human motion. W…

Video Motion Capture from the Part Confidence Maps of Multi-Camera Images by Spatiotemporal Filtering Using the Human Skeletal Model

2019-12-09 · Takuya Ohashi, Yosuke Ikegami, Kazuki Yamamoto, Wataru Takano 외

This paper discusses video motion capture, namely, 3D reconstruction of human motion from multi-camera images. After the Part Confidence Maps are computed from each camera image, the proposed spatiotemporal filter is app…

3D ReconstructionPosition

Combining Spatio-Temporal Appearance Descriptors and Optical Flow for Human Action Recognition in Video Data

2013-10-01 · Karla Brkić, Srđan Rašić, Axel Pinz, Siniša Šegvić 외

This paper proposes combining spatio-temporal appearance (STA) descriptors with optical flow for human action recognition. The STA descriptors are local histogram-based descriptors of space-time, suitable for building a …

Action RecognitionOptical Flow EstimationTemporal Action Localization

Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding

2025-11-12 · Jingtian Ma, Jingyuan Wang, Wayne Xin Zhao, Guoping Liu 외 arxiv

Nowadays, navigation and ride-sharing apps have collected numerous images with spatio-temporal data. A core technology for analyzing such images, associated with spatiotemporal information, is Traffic Scene Understanding…

Scene UnderstandingFew-Shot Learning