paper-with-me

Papers

Depth Pooling Based Large-scale 3D Action Recognition with Convolutional Neural Networks

2018-03-17 · Pichao Wang, Wanqing Li, Zhimin Gao, Chang Tang, Philip Ogunbona

This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as Dynamic Depth Images (DDI), Dynamic Depth Normal Images (DDNI) and Dynamic Depth Motion Normal Images (DDMNI), for both isolated and continuous action recognition. These dynamic images are constructed from a segmented sequence of depth maps using hierarchical bidirectional rank pooling to effectively capture the spatial-temporal information. Specifically, DDI exploits the dynamics of postures over time and DDNI and DDMNI exploit the 3D structural information captured by depth maps. Upon the proposed representations, a ConvNet based method is developed for action recognition. The image-based representations enable us to fine-tune the existing Convolutional Neural Network (ConvNet) models trained on image data without training a large number of parameters from scratch. The proposed method achieved the state-of-art results on three large datasets, namely, the Large-scale Continuous Gesture Recognition Dataset (means Jaccard index 0.4109), the Large-scale Isolated Gesture Recognition Dataset (59.21%), and the NTU RGB+D Dataset (87.08% cross-subject and 84.22% cross-view) even though only the depth modality was used.

📄 PDF Abstract BibTeX arXiv:1804.01194

Code (0)

등록된 구현이 없습니다.

Tasks

3D Action RecognitionAction RecognitionGesture RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Large-scale Isolated Gesture Recognition Using Convolutional Neural Networks

2017-01-07 · Pichao Wang, Wanqing Li, Song Liu, Zhimin Gao 외

This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as Dynamic Depth Images (DDI), Dynamic Depth Normal Images (DDNI) and Dynamic Depth Motion Normal Image…

General ClassificationGesture Recognition

Recurrent Scene Parsing with Perspective Understanding in the Loop

2017-05-20 · CVPR 2018 6 · Shu Kong, Charless Fowlkes

Objects may appear at arbitrary scales in perspective images of a scene, posing a challenge for recognition systems that process images at a fixed resolution. We propose a depth-aware gating module that adaptively select…

Depth EstimationMonocular Depth EstimationScene ParsingSegmentation+1

Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition

2016-08-15 · Yi Zhu, Shawn Newsam

This paper performs the first investigation into depth for large-scale human action recognition in video where the depth cues are estimated from the videos themselves. We develop a new framework called depth2action and e…

Action RecognitionTemporal Action Localization

GestFormer: Multiscale Wavelet Pooling Transformer Network for Dynamic Hand Gesture Recognition

2024-05-18 · Mallika Garg, Debashis Ghosh, Pyari Mohan Pradhan

Transformer model have achieved state-of-the-art results in many applications like NLP, classification, etc. But their exploration in gesture recognition task is still limited. So, we propose a novel GestFormer architect…

Gesture RecognitionHand Gesture RecognitionHand-Gesture RecognitionOptical Flow Estimation

3DV: 3D Dynamic Voxel for Action Recognition in Depth Video

2020-05-12 · CVPR 2020 6 · Yancheng Wang, Yang Xiao, Fu Xiong, Wenxiang Jiang 외

To facilitate depth-based 3D action recognition, 3D dynamic voxel (3DV) is proposed as a novel 3D motion representation. With 3D space voxelization, the key idea of 3DV is to encode 3D motion information within depth vid…

3D Action RecognitionAction Recognition