ConvNet Architecture Search for Spatiotemporal Feature Learning
Learning image representations with ConvNets by pre-training on ImageNet has proven useful across many visual understanding tasks including object detection, semantic segmentation, and image captioning. Although any image representation can be applied to video frames, a dedicated spatiotemporal representation is still vital in order to incorporate motion patterns that cannot be captured by appearance based models alone. This paper presents an empirical ConvNet architecture search for spatiotemporal feature learning, culminating in a deep 3-dimensional (3D) Residual ConvNet. Our proposed architecture outperforms C3D by a good margin on Sports-1M, UCF101, HMDB51, THUMOS14, and ASLAN while being 2 times faster at inference time, 2 times smaller in model size, and having a more compact representation.
Code (1)
Tasks
Action ClassificationAction RecognitionImage CaptioningNeural Architecture Searchobject-detectionObject DetectionSemantic SegmentationSimilar Papers 제목 키워드 기반
Learning Spatiotemporal Features with 3D Convolutional Networks
We propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset. Our findings are three-fold…
Action RecognitionAction Recognition In VideosDynamic Facial Expression RecognitionSpatiotemporal Residual Networks for Video Action Recognition
Two-stream Convolutional Networks (ConvNets) have shown strong performance for human action recognition in videos. Recently, Residual Networks (ResNets) have arisen as a new technique to train extremely deep architecture…
Action RecognitionAction Recognition In VideosTemporal Action LocalizationTS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition
Recent two-stream deep Convolutional Neural Networks (ConvNets) have made significant progress in recognizing human actions in videos. Despite their success, methods extending the basic two-stream ConvNet have not system…
Action ClassificationAction RecognitionActivity RecognitionTemporal Action Localization+2Spatio-Temporal Graph Convolutional Neural Networks for Physics-Aware Grid Learning Algorithms
This paper proposes a model-free Volt-VAR control (VVC) algorithm via the spatio-temporal graph ConvNet-based deep reinforcement learning (STGCN-DRL) framework, whose goal is to control smart inverters in an unbalanced d…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Convolutional Two-Stream Network Fusion for Video Action Recognition
Recent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information. We study a number of ways …
Action RecognitionAction Recognition In VideosTemporal Action LocalizationVocal Bursts Valence Prediction