Spatiotemporal Deformable Part Models for Action Detection
Deformable part models have achieved impressive performance for object detection, even on difficult image datasets. This paper explores the generalization of deformable part models from 2D images to 3D spatiotemporal volumes to better study their effectiveness for action detection in video. Actions are treated as spatiotemporal patterns and a deformable part model is generated for each action from a collection of examples. For each action model, the most discriminative 3D subvolumes are automatically selected as parts and the spatiotemporal relations between their locations are learned. By focusing on the most distinctive parts of each action, our models adapt to intra-class variation and show robustness to clutter. Extensive experiments on several video datasets demonstrate the strength of spatiotemporal DPMs for classifying and localizing actions.
Code (0)
등록된 구현이 없습니다.
Tasks
Action Detectionobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
Spatiotemporal Deformable Scene Graphs for Complex Activity Detection
Long-term complex activity recognition and localisation can be crucial for decision making in autonomous systems such as smart cars and surgical robots. Here we address the problem via a novel deformable, spatiotemporal …
Action DetectionActivity DetectionActivity RecognitionAutonomous Driving+1Spatiotemporal Event Graphs for Dynamic Scene Understanding
Dynamic scene understanding is the ability of a computer system to interpret and make sense of the visual information present in a video of a real-world scene. In this thesis, we present a series of frameworks for dynami…
Action DetectionActivity DetectionAutonomous DrivingContinual Learning+3Compositional Structure Learning for Action Understanding
The focus of the action understanding literature has predominately been classification, how- ever, there are many applications demanding richer action understanding such as mobile robotics and video search, with solution…
Action DetectionAction UnderstandingGeneral ClassificationCross-Modal Learning with 3D Deformable Attention for Action Recognition
An important challenge in vision-based action recognition is the embedding of spatiotemporal features with two or more heterogeneous modalities into a single feature. In this study, we propose a new 3D deformable transfo…
Action RecognitionUnconstrained Monocular 3D Human Pose Estimation by Action Detection and Cross-Modality Regression Forest
This work addresses the challenging problem of unconstrained 3D human pose estimation (HPE) from a novel perspective. Existing approaches struggle to operate in realistic applications, mainly due to their scene-dependent…
2D Pose Estimation3D Human Pose EstimationAction DetectionMonocular 3D Human Pose Estimation+3