Action Understanding with Multiple Classes of Actors
Despite the rapid progress, existing works on action understanding focus strictly on one type of action agent, which we call actor---a human adult, ignoring the diversity of actions performed by other actors. To overcome this narrow viewpoint, our paper marks the first effort in the computer vision community to jointly consider algorithmic understanding of various types of actors undergoing various actions. To begin with, we collect a large annotated Actor-Action Dataset (A2D) that consists of 3782 short videos and 31 temporally untrimmed long videos. We formulate the general actor-action understanding problem and instantiate it at various granularities: video-level single- and multiple-label actor-action recognition, and pixel-level actor-action segmentation. We propose and examine a comprehensive set of graphical models that consider the various types of interplay among actors and actions. Our findings have led us to conclusive evidence that the joint modeling of actor and action improves performance over modeling each of them independently, and further improvement can be obtained by considering the multi-scale natural in video understanding. Hence, our paper concludes the argument of the value of explicit consideration of various actors in comprehensive action understanding and provides a dataset and a benchmark for later works exploring this new problem.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction SegmentationAction UnderstandingDiversityTemporal Action LocalizationVideo UnderstandingSimilar Papers 제목 키워드 기반
Can Humans Fly? Action Understanding With Multiple Classes of Actors
Can humans fly? Emphatically no. Can cars eat? Again, absolutely not. Yet, these absurd inferences result from the current disregard for particular types of actors in action understanding. There is no work we know of on …
Action RecognitionAction UnderstandingSemantic SegmentationTemporal Action LocalizationRepresentation Learning on Visual-Symbolic Graphs for Video Understanding
Events in natural videos typically arise from spatio-temporal interactions between actors and objects and involve multiple co-occurring activities and object classes. To capture this rich visual and semantic context, we …
Action ClassificationAction DetectionAction LocalizationAction Segmentation+5End-to-End Joint Semantic Segmentation of Actors and Actions in Video
Traditional video understanding tasks include human action recognition and actor/object semantic segmentation. However, the combined task of providing semantic segmentation for different actor classes simultaneously with…
Action RecognitionSegmentationSemantic SegmentationTemporal Action Localization+2On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices
While learning with limited labelled data can improve performance when the labels are lacking, it is also sensitive to the effects of uncontrolled randomness introduced by so-called randomness factors (e.g., varying orde…
In-Context LearningMeta-LearningSensitivitytext-classification+1Object Level Visual Reasoning in Videos
Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The ne…
Activity RecognitionHuman Activity RecognitionObjectobject-detection+2