paper-with-me

홈 › Papers

I3D-LSTM: A New Model for Human Action Recognition

2019-08-09 · IOP Conf. Ser.: Mater. Sci. Eng. 569 032035 2019 8 · Xianyuan Wang, Zhenjiang Miao, Ruyi Zhang, Shanshan Hao

Action recognition has already been a heated research topic recently, which attempts to classify different human actions in videos. The current main-stream methods generally utilize ImageNet-pretrained model as features extractor, however it's not the optimal choice to pretrain a model for classifying videos on a huge still image dataset. What's more, very few works notice that 3D convolution neural network(3D CNN) is better for low-level spatial-temporal features extraction while recurrent neural network(RNN) is better for modelling high-level temporal feature sequences. Consequently, a novel model is proposed in our work to address the two problems mentioned above. First, we pretrain 3D CNN model on huge video action recognition dataset Kinetics to improve generality of the model. And then long short term memory(LSTM) is introduced to model the high-level temporal features produced by the Kinetics-pretrained 3D CNN model. Our experiments results show that the Kinetics-pretrained model can generally outperform ImageNet-pretrained model. And our proposed network finally achieve leading performance on UCF-101 dataset.

📄 PDF Abstract BibTeX

Code (1)

Mind23-2/MindCode-101/tree/main/I3D mindspore

Tasks

Action RecognitionTemporal Action Localization

Methods 이 논문이 사용한 방법론

3D Convolution A 3D Convolution is a type of convolution where the kernel slides in 3 dimensions as opposed to 2 dimensions with 2D…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Co-occurrence Feature Learning for Skeleton based Action Recognition using Regularized Deep LSTM Networks

2016-03-24 · Wentao Zhu, Cuiling Lan, Junliang Xing, Wen-Jun Zeng 외

Skeleton based action recognition distinguishes human actions using the trajectories of skeleton joints, which provide a very good representation for describing actions. Considering that recurrent neural networks (RNNs) …

Action RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Regularizing Long Short Term Memory With 3D Human-Skeleton Sequences for Action Recognition

2016-06-01 · CVPR 2016 6 · Behrooz Mahasseni, Sinisa Todorovic

This paper argues that large-scale action recognition in video can be greatly improved by providing an additional modality in training data -- namely, 3D human-skeleton sequences -- aimed at complementing poorly represen…

Action RecognitionTemporal Action Localization

Skeleton-Based Human Action Recognition with Global Context-Aware Attention LSTM Networks

2017-07-18 · Jun Liu, Gang Wang, Ling-Yu Duan, Kamila Abdiyeva 외

Human action recognition in 3D skeleton sequences has attracted a lot of research attention. Recently, Long Short-Term Memory (LSTM) networks have shown promising performance in this task due to their strengths in modeli…

Action RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Speech Emotion Recognition Based on CNN+LSTM Model

2021-10-01 · ROCLING 2021 10 · Wei Mou, Pei-Hsuan Shen, Chu-Yun Chu, Yu-Cheng Chiu 외

Due to the popularity of intelligent dialogue assistant services, speech emotion recognition has become more and more important. In the communication between humans and machines, emotion recognition and emotion analysis …

Emotion RecognitionmodelSpeech Emotion Recognition

Hierarchical Long Short-Term Concurrent Memory for Human Interaction Recognition

2018-11-01 · Xiangbo Shu, Jinhui Tang, Guo-Jun Qi, Wei Liu 외

In this paper, we aim to address the problem of human interaction recognition in videos by exploring the long-term inter-related dynamics among multiple persons. Recently, Long Short-Term Memory (LSTM) has become a popul…

Action RecognitionHuman Interaction RecognitionTemporal Action Localization