Compositional Structure Learning for Action Understanding
The focus of the action understanding literature has predominately been classification, how- ever, there are many applications demanding richer action understanding such as mobile robotics and video search, with solutions to classification, localization and detection. In this paper, we propose a compositional model that leverages a new mid-level representation called compositional trajectories and a locally articulated spatiotemporal deformable parts model (LALSDPM) for fully action understanding. Our methods is advantageous in capturing the variable structure of dynamic human activity over a long range. First, the compositional trajectories capture long-ranging, frequently co-occurring groups of trajectories in space time and represent them in discriminative hierarchies, where human motion is largely separated from camera motion; second, LASTDPM learns a structured model with multi-layer deformable parts to capture multiple levels of articulated motion. We implement our methods and demonstrate state of the art performance on all three problems: action detection, localization, and recognition.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionAction UnderstandingGeneral ClassificationSimilar Papers 제목 키워드 기반
MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic pat…
Action Quality AssessmentAction UnderstandingUnderstanding Human Actions through the Lens of Executable Models
Human-centred systems require an understanding of human actions in the physical world. Temporally extended sequences of actions are intentional and structured, yet existing methods for recognising what actions are perfor…
Action SegmentationAnomaly DetectionTemporal Modular Networks for Retrieving Complex Compositional Activities in Videos
A major challenge in computer vision is scaling activity understanding to the long tail of complex activities without requiring collecting large quantities of data for new actions. The task of video retrieval using natur…
RetrievalVideo RetrievalObject-centric Binding in Contrastive Language-Image Pretraining
Recent advances in vision language models (VLM) have been driven by contrastive models such as CLIP, which learn to associate visual information with their corresponding text descriptions. However, these models have limi…
Image-text matchingObjectText MatchingWeak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
Vision-language (VL) models often exhibit a limited understanding of complex expressions of visual objects (e.g., attributes, shapes, and their relations), given complex and diverse language queries. Traditional approach…
Contrastive Learningobject-detectionObject DetectionSynthetic Data Generation