Spatio-Temporal Vector of Locally Max Pooled Features for Action Recognition in Videos
We introduce Spatio-Temporal Vector of Locally Max Pooled Features (ST-VLMPF), a super vector-based encoding method specifically designed for local deep features encoding. The proposed method addresses an important problem of video understanding: how to build a video representation that incorporates the CNN features over the entire video. Feature assignment is carried out at two levels, by using the similarity and spatio-temporal information. For each assignment we build a specific encoding, focused on the nature of deep features, with the goal to capture the highest feature responses from the highest neuron activation of the network. Our ST-VLMPF clearly provides a more reliable video representation than some of the most widely used and powerful encoding approaches (Improved Fisher Vectors and Vector of Locally Aggregated Descriptors), while maintaining a low computational complexity. We conduct experiments on three action recognition datasets: HMDB51, UCF50 and UCF101. Our pipeline obtains state-of-the-art results.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction Recognition In VideosTemporal Action LocalizationVideo UnderstandingSimilar Papers 제목 키워드 기반
Multiresolution Match Kernels for Gesture Video Classification
The emergence of depth imaging technologies like the Microsoft Kinect has renewed interest in computational methods for gesture classification based on videos. For several years now, researchers have used the Bag-of-Feat…
ClassificationGeneral ClassificationVideo ClassificationLearning Motion in Feature Space: Locally-Consistent Deformable Convolution Networks for Fine-Grained Action Detection
Fine-grained action detection is an important task with numerous applications in robotics and human-computer interaction. Existing methods typically utilize a two-stage approach including extraction of local spatio-tempo…
Action DetectionFine-Grained Action DetectionOptical Flow EstimationAction Recognition with Trajectory-Pooled Deep-Convolutional Descriptors
Visual features are of vital importance for human action understanding in videos. This paper presents a new video representation, called trajectory-pooled deep-convolutional descriptor (TDD), which shares the merits of b…
Action RecognitionAction UnderstandingActivity Recognition In VideosTemporal Action LocalizationLearning Spatiotemporal Features of Ride-sourcing Services with Fusion Convolutional Network
To collectively forecast the demand for ride-sourcing services in all regions of a city, the deep learning approaches have been applied with commendable results. However, the local statistical differences throughout the …
Demand ForecastingSpatio-temporal Tubelet Feature Aggregation and Object Linking in Videos
This paper addresses the problem of how to exploit spatio-temporal information available in videos to improve the object detection precision. We propose a two stage object detector called FANet based on short-term spatio…
Objectobject-detectionObject DetectionObject Localization+1