paper-with-me

홈 › Papers

Learning Gating ConvNet for Two-Stream based Methods in Action Recognition

2017-09-12 · Jiagang Zhu, Wei Zou, Zheng Zhu

For the two-stream style methods in action recognition, fusing the two streams' predictions is always by the weighted averaging scheme. This fusion method with fixed weights lacks of pertinence to different action videos and always needs trial and error on the validation set. In order to enhance the adaptability of two-stream ConvNets and improve its performance, an end-to-end trainable gated fusion method, namely gating ConvNet, for the two-stream ConvNets is proposed in this paper based on the MoE (Mixture of Experts) theory. The gating ConvNet takes the combination of feature maps from the same layer of the spatial and the temporal nets as input and adopts ReLU (Rectified Linear Unit) as the gating output activation function. To reduce the over-fitting of gating ConvNet caused by the redundancy of parameters, a new multi-task learning method is designed, which jointly learns the gating fusion weights for the two streams and learns the gating ConvNet for action classification. With our gated fusion method and multi-task learning approach, a high accuracy of 94.5% is achieved on the dataset UCF101.

📄 PDF Abstract BibTeX arXiv:1709.03655

Code (1)

zhujiagang/gating-ConvNet-code 공식 구현

Tasks

Action ClassificationAction RecognitionMixture-of-ExpertsMulti-Task LearningTemporal Action LocalizationVocal Bursts Valence Prediction

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Spatiotemporal Multiplier Networks for Video Action Recognition

2017-07-01 · CVPR 2017 7 · Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes

This paper presents a general ConvNet architecture for video action recognition based on multiplicative interactions of spacetime features. Our model combines the appearance and motion pathways of a two-stream architectu…

Action RecognitionGeneral ClassificationTemporal Action Localization

Towards Good Practices for Very Deep Two-Stream ConvNets

2015-07-08 · Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao

Deep convolutional networks have achieved great success for object recognition in still images. However, for action recognition in videos, the improvement of deep convolutional networks is not so evident. We argue that t…

Action RecognitionAction Recognition In VideosComputational EfficiencyData Augmentation+3

Semi-Coupled Two-Stream Fusion ConvNets for Action Recognition at Extremely Low Resolutions

2016-10-12 · Jiawei Chen, Jonathan Wu, Janusz Konrad, Prakash Ishwar

Deep convolutional neural networks (ConvNets) have been recently shown to attain state-of-the-art performance for action recognition on standard-resolution videos. However, less attention has been paid to recognition per…

Action RecognitionTemporal Action Localization

Cross-Enhancement Transform Two-Stream 3D ConvNets for Action Recognition

2019-08-19 · Dong Cao, Lisha Xu, Dong-dong Zhang

Action recognition is an important research topic in computer vision. It is the basic work for visual understanding and has been applied in many fields. Since human actions can vary in different environments, it is diffi…

Action RecognitionAutonomous DrivingAutonomous VehiclesOptical Flow Estimation+2

Low-Latency Human Action Recognition with Weighted Multi-Region Convolutional Neural Network

2018-05-08 · Yunfeng Wang, Wengang Zhou, Qilin Zhang, Xiaotian Zhu 외

Spatio-temporal contexts are crucial in understanding human actions in videos. Recent state-of-the-art Convolutional Neural Network (ConvNet) based action recognition systems frequently involve 3D spatio-temporal ConvNet…

Action RecognitionChunkingOptical Flow EstimationTemporal Action Localization