paper-with-me

홈 › Papers

Can Humans Fly? Action Understanding With Multiple Classes of Actors

2015-06-01 · CVPR 2015 6 · Chenliang Xu, Shao-Hang Hsieh, Caiming Xiong, Jason J. Corso

Can humans fly? Emphatically no. Can cars eat? Again, absolutely not. Yet, these absurd inferences result from the current disregard for particular types of actors in action understanding. There is no work we know of on simultaneously inferring actors and actions in the video, not to mention a dataset to experiment with. Our paper hence marks the first effort in the computer vision community to jointly consider various types of actors undergoing various actions. To start with the problem, we collect a dataset of 3782 videos from YouTube and label both pixel-level actors and actions in each video. We formulate the general actor-action understanding problem and instantiate it at various granularities: both video-level single- and multiple-label actor-action recognition and pixel-level actor-action semantic segmentation. Our experiments demonstrate that inference jointly over actors and actions outperforms inference independently over them, and hence concludes our argument of the value of explicit consideration of various actors in comprehensive action understanding.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionAction UnderstandingSemantic SegmentationTemporal Action Localization

Similar Papers 제목 키워드 기반

Action Understanding with Multiple Classes of Actors

2017-04-27 · Chenliang Xu, Caiming Xiong, Jason J. Corso

Despite the rapid progress, existing works on action understanding focus strictly on one type of action agent, which we call actor---a human adult, ignoring the diversity of actions performed by other actors. To overcome…

Action RecognitionAction SegmentationAction UnderstandingDiversity+2

MOMA: Multi-Object Multi-Actor Activity Parsing

2021-12-01 · NeurIPS 2021 12 · Zelun Luo, Wanze Xie, Siddharth Kapoor, Yiyun Liang 외

Complex activities often involve multiple humans utilizing different objects to complete actions (e.g., in healthcare settings, physicians, nurses, and patients interact with each other and various medical devices). Reco…

Object

Actor-agnostic Multi-label Action Recognition with Multi-modal Query

2023-07-20 · Anindya Mondal, Sauradip Nag, Joaquin M Prada, Xiatian Zhu 외

Existing action recognition methods are typically actor-specific due to the intrinsic topological and apparent differences among the actors. This requires actor-specific pose estimation (e.g., humans vs. animals), leadin…

Action ClassificationAction RecognitionAction Recognition In VideosAction Recognition on HMDB-51+2

Representation Learning on Visual-Symbolic Graphs for Video Understanding

2019-05-17 · ECCV 2020 8 · Effrosyni Mavroudi, Benjamín Béjar Haro, René Vidal

Events in natural videos typically arise from spatio-temporal interactions between actors and objects and involve multiple co-occurring activities and object classes. To capture this rich visual and semantic context, we …

Action ClassificationAction DetectionAction LocalizationAction Segmentation+5

SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos

2024-04-06 · CVPR 2024 1 · Tao Wu, Runyu He, Gangshan Wu, LiMin Wang

Video-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation detection datasets have two main limitat…

Graph GenerationRelationScene Graph GenerationVideo scene graph generation+2