Finding Action Tubes
We address the problem of action detection in videos. Driven by the latest progress in object detection from 2D images, we build action models using rich feature hierarchies derived from shape and kinematic cues. We incorporate appearance and motion in two ways. First, starting from image region proposals we select those that are motion salient and thus are more likely to contain the action. This leads to a significant reduction in the number of regions being processed and allows for faster computations. Second, we extract spatio-temporal feature representations to build strong classifiers using Convolutional Neural Networks. We link our predictions to produce detections consistent in time, which we call action tubes. We show that our approach outperforms other techniques in the task of action detection.
Code (1)
Tasks
Action Detectionobject-detectionObject DetectionSkeleton Based Action RecognitionSimilar Papers 제목 키워드 기반
Human Action Localization with Sparse Spatial Supervision
We introduce an approach for spatio-temporal human action localization using sparse spatial supervision. Our method leverages the large amount of annotated humans available today and extracts human tubes by combining a s…
Action LocalizationDiversityDiscovering Spatio-Temporal Action Tubes
In this paper, we address the challenging problem of spatial and temporal action detection in videos. We first develop an effective approach to localize frame-level action regions through integrating static and kinematic…
Action DetectionSystem Level Disturbance Reachable Sets and their Application to Tube-based MPC
Tube-based model predictive control (MPC) methods leverage tubes to bound deviations from a nominal trajectory due to uncertainties in order to ensure constraint satisfaction. This paper presents a novel tube-based MPC f…
Model Predictive ControlPredicting Action Tubes
In this work, we present a method to predict an entire `action tube' (a set of temporally linked bounding boxes) in a trimmed video just by observing a smaller subset of it. Predicting where an action is going to take pl…
Action ClassificationAction DetectionAutonomous DrivingDeformable Tube Network for Action Detection in Videos
We address the problem of spatio-temporal action detection in videos. Existing methods commonly either ignore temporal context in action recognition and localization, or lack the modelling of flexible shapes of action tu…
Action DetectionAction Recognition