paper-with-me

홈 › Papers

What and Where: Modeling Skeletons from Semantic and Spatial Perspectives for Action Recognition

2020-04-07 · Lei Shi, Yifan Zhang, Jian Cheng, Hanqing Lu

Skeleton data, which consists of only the 2D/3D coordinates of the human joints, has been widely studied for human action recognition. Existing methods take the semantics as prior knowledge to group human joints and draw correlations according to their spatial locations, which we call the semantic perspective for skeleton modeling. In this paper, in contrast to previous approaches, we propose to model skeletons from a novel spatial perspective, from which the model takes the spatial location as prior knowledge to group human joints and mines the discriminative patterns of local areas in a hierarchical manner. The two perspectives are orthogonal and complementary to each other; and by fusing them in a unified framework, our method achieves a more comprehensive understanding of the skeleton data. Besides, we customized two networks for the two perspectives. From the semantic perspective, we propose a Transformer-like network that is expert in modeling joint correlations, and present three effective techniques to adapt it for skeleton data. From the spatial perspective, we transform the skeleton data into the sparse format for efficient feature extraction and present two types of sparse convolutional networks for sparse skeleton modeling. Extensive experiments are conducted on three challenging datasets for skeleton-based human action/gesture recognition, namely, NTU-60, NTU-120 and SHREC, where our method achieves state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2004.03259

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionGesture RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition

2018-01-23 · Sijie Yan, Yuanjun Xiong, Dahua Lin

Dynamics of human body skeletons convey significant information for human action recognition. Conventional approaches for modeling skeletons usually rely on hand-crafted parts or traversal rules, thus resulting in limite…

3D Human Pose EstimationAction RecognitionMultimodal Activity RecognitionSkeleton Based Action Recognition+1

SceneBind: Binding What and Where Across Vision, Audio and Language

2026-07-16 · Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu 외 arxiv

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i…

Visual Localization

Learning Rich Features for Gait Recognition by Integrating Skeletons and Silhouettes

2021-10-26 · Yunjie Peng, Kang Ma, Yang Zhang, Zhiqiang He

Gait recognition captures gait patterns from the walking sequence of an individual for identification. Most existing gait recognition methods learn features from silhouettes or skeletons for the robustness to clothing, c…

Gait IdentificationGait Recognition

Modeling Temporal Dynamics and Spatial Configurations of Actions Using Two-Stream Recurrent Neural Networks

2017-04-09 · CVPR 2017 7 · Hongsong Wang, Liang Wang

Recently, skeleton based action recognition gains more popularity due to cost-effective depth sensors coupled with real-time skeleton estimation algorithms. Traditional approaches based on handcrafted features are limite…

3D Action RecognitionAction RecognitionData AugmentationSkeleton Based Action Recognition+1

Where, What, Why: Towards Explainable Driver Attention Prediction

2025-06-29 · Yuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin 외

Modeling task-driven attention in driving is a fundamental challenge for both autonomous vehicles and cognitive science. Existing methods primarily predict where drivers look by generating spatial heatmaps, but fail to c…

Autonomous DrivingAutonomous VehiclesDriver Attention MonitoringLarge Language Model+2