paper-with-me

Papers

ConvNet Architecture Search for Spatiotemporal Feature Learning

2017-08-16 · Du Tran, Jamie Ray, Zheng Shou, Shih-Fu Chang, Manohar Paluri

Learning image representations with ConvNets by pre-training on ImageNet has proven useful across many visual understanding tasks including object detection, semantic segmentation, and image captioning. Although any image representation can be applied to video frames, a dedicated spatiotemporal representation is still vital in order to incorporate motion patterns that cannot be captured by appearance based models alone. This paper presents an empirical ConvNet architecture search for spatiotemporal feature learning, culminating in a deep 3-dimensional (3D) Residual ConvNet. Our proposed architecture outperforms C3D by a good margin on Sports-1M, UCF101, HMDB51, THUMOS14, and ASLAN while being 2 times faster at inference time, 2 times smaller in model size, and having a more compact representation.

📄 PDF Abstract BibTeX arXiv:1708.05038

Code (1)

farazahmeds/Classification-of-brain-tumor-using-Spatiotemporal-models pytorch

Tasks

Action ClassificationAction RecognitionImage CaptioningNeural Architecture Searchobject-detectionObject DetectionSemantic Segmentation

Similar Papers 제목 키워드 기반

Learning Spatiotemporal Features with 3D Convolutional Networks

2014-12-02 · ICCV 2015 12 · Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani 외

We propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset. Our findings are three-fold…

Action RecognitionAction Recognition In VideosDynamic Facial Expression Recognition

Spatiotemporal Residual Networks for Video Action Recognition

2016-11-07 · NeurIPS 2016 12 · Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes

Two-stream Convolutional Networks (ConvNets) have shown strong performance for human action recognition in videos. Recently, Residual Networks (ResNets) have arisen as a new technique to train extremely deep architecture…

Action RecognitionAction Recognition In VideosTemporal Action Localization

TS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition

2017-03-30 · Chih-Yao Ma, Min-Hung Chen, Zsolt Kira, Ghassan AlRegib

Recent two-stream deep Convolutional Neural Networks (ConvNets) have made significant progress in recognizing human actions in videos. Despite their success, methods extending the basic two-stream ConvNet have not system…

Action ClassificationAction RecognitionActivity RecognitionTemporal Action Localization+2

Spatio-Temporal Graph Convolutional Neural Networks for Physics-Aware Grid Learning Algorithms

2022-03-31 · Tong Wu, Ignacio Losada Carreno, Anna Scaglione, Daniel Arnold

This paper proposes a model-free Volt-VAR control (VVC) algorithm via the spatio-temporal graph ConvNet-based deep reinforcement learning (STGCN-DRL) framework, whose goal is to control smart inverters in an unbalanced d…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Convolutional Two-Stream Network Fusion for Video Action Recognition

2016-04-22 · CVPR 2016 6 · Christoph Feichtenhofer, Axel Pinz, Andrew Zisserman

Recent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information. We study a number of ways …

Action RecognitionAction Recognition In VideosTemporal Action LocalizationVocal Bursts Valence Prediction