Learning Gating ConvNet for Two-Stream based Methods in Action Recognition
For the two-stream style methods in action recognition, fusing the two streams' predictions is always by the weighted averaging scheme. This fusion method with fixed weights lacks of pertinence to different action videos and always needs trial and error on the validation set. In order to enhance the adaptability of two-stream ConvNets and improve its performance, an end-to-end trainable gated fusion method, namely gating ConvNet, for the two-stream ConvNets is proposed in this paper based on the MoE (Mixture of Experts) theory. The gating ConvNet takes the combination of feature maps from the same layer of the spatial and the temporal nets as input and adopts ReLU (Rectified Linear Unit) as the gating output activation function. To reduce the over-fitting of gating ConvNet caused by the redundancy of parameters, a new multi-task learning method is designed, which jointly learns the gating fusion weights for the two streams and learns the gating ConvNet for action classification. With our gated fusion method and multi-task learning approach, a high accuracy of 94.5% is achieved on the dataset UCF101.
Code (1)
Tasks
Action ClassificationAction RecognitionMixture-of-ExpertsMulti-Task LearningTemporal Action LocalizationVocal Bursts Valence PredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Spatiotemporal Multiplier Networks for Video Action Recognition
This paper presents a general ConvNet architecture for video action recognition based on multiplicative interactions of spacetime features. Our model combines the appearance and motion pathways of a two-stream architectu…
Action RecognitionGeneral ClassificationTemporal Action LocalizationTowards Good Practices for Very Deep Two-Stream ConvNets
Deep convolutional networks have achieved great success for object recognition in still images. However, for action recognition in videos, the improvement of deep convolutional networks is not so evident. We argue that t…
Action RecognitionAction Recognition In VideosComputational EfficiencyData Augmentation+3Semi-Coupled Two-Stream Fusion ConvNets for Action Recognition at Extremely Low Resolutions
Deep convolutional neural networks (ConvNets) have been recently shown to attain state-of-the-art performance for action recognition on standard-resolution videos. However, less attention has been paid to recognition per…
Action RecognitionTemporal Action LocalizationCross-Enhancement Transform Two-Stream 3D ConvNets for Action Recognition
Action recognition is an important research topic in computer vision. It is the basic work for visual understanding and has been applied in many fields. Since human actions can vary in different environments, it is diffi…
Action RecognitionAutonomous DrivingAutonomous VehiclesOptical Flow Estimation+2Low-Latency Human Action Recognition with Weighted Multi-Region Convolutional Neural Network
Spatio-temporal contexts are crucial in understanding human actions in videos. Recent state-of-the-art Convolutional Neural Network (ConvNet) based action recognition systems frequently involve 3D spatio-temporal ConvNet…
Action RecognitionChunkingOptical Flow EstimationTemporal Action Localization