Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos
Current state-of-the-art human action recognition is focused on the classification of temporally trimmed videos in which only one action occurs per frame. In this work we address the problem of action localisation and instance segmentation in which multiple concurrent actions of the same class may be segmented out of an image sequence. We cast the action tube extraction as an energy maximisation problem in which configurations of region proposals in each frame are assigned a cost and the best action tubes are selected via two passes of dynamic programming. One pass associates region proposals in space and time for each action category, and another pass is used to solve for the tube's temporal extent and to enforce a smooth label sequence through the video. In addition, by taking advantage of recent work on action foreground-background segmentation, we are able to associate each tube with class-specific segmentations. We demonstrate the performance of our algorithm on the challenging LIRIS-HARL dataset and achieve a new state-of-the-art result which is 14.3 times better than previous methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionInstance SegmentationSemantic SegmentationTemporal Action LocalizationSimilar Papers 제목 키워드 기반
End-to-End Spatio-Temporal Action Localisation with Video Transformers
The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end, purely-transformer based model that directly ingests an input…
Action DetectionAction RecognitionSpatio-Temporal Action LocalizationEnhancing super-resolution ultrasound localisation through multi-frame deconvolution exploiting spatiotemporal coherence
Super-resolution ultrasound imaging through microbubble (MB) localisation and tracking, also known as ultrasound localisation microscopy, allows non-invasive sub-diffraction resolution imaging of microvasculature in anim…
DenoisingSuper-ResolutionOnline Real-time Multiple Spatiotemporal Action Localisation and Prediction
We present a deep-learning framework for real-time multiple spatio-temporal (S/T) action localisation, classification and early prediction. Current state-of-the-art approaches work offline and are too slow to be useful i…
Early Action PredictionPredictionSpatio-Temporal Instance Learning: Action Tubes from Class Supervision
The goal of this work is spatio-temporal action localization in videos, using only the supervision from video-level class labels. The state-of-the-art casts this weakly-supervised action localization regime as a Multiple…
Action LocalizationMultiple Instance LearningRerankingSpatio-Temporal Action Localization+2Human Action Localization with Sparse Spatial Supervision
We introduce an approach for spatio-temporal human action localization using sparse spatial supervision. Our method leverages the large amount of annotated humans available today and extracts human tubes by combining a s…
Action LocalizationDiversity