Actor-centered Representations for Action Localization in Streaming Videos
Event perception tasks such as recognizing and localizing actions in streaming videos are essential for scaling to real-world application contexts. We tackle the problem of learning actor-centered representations through the notion of continual hierarchical predictive learning to localize actions in streaming videos without the need for training labels and outlines for the objects in the video. We propose a framework driven by the notion of hierarchical predictive learning to construct actor-centered features by attention-based contextualization. The key idea is that predictable features or objects do not attract attention and hence do not contribute to the action of interest. Experiments on three benchmark datasets show that the approach can learn robust representations for localizing actions using only one epoch of training, i.e., a single pass through the streaming video. We show that the proposed approach outperforms unsupervised and weakly supervised baselines while offering competitive performance to fully supervised approaches. Additionally, we extend the model to multi-actor settings to recognize group activities while localizing the multiple, plausible actors. We also show that it generalizes to out-of-domain data with limited performance degradation.
Code (0)
등록된 구현이 없습니다.
Tasks
Action LocalizationSimilar Papers 제목 키워드 기반
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
This work addresses the problem of Social Activity Recognition (SAR), a critical component in real-world tasks like surveillance and assistive robotics. Unlike traditional event understanding approaches, SAR necessitates…
Action LocalizationActivity RecognitionGroup Activity RecognitionRelational ReasoningMulti-HMR 2: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking
Most advances in human mesh recovery (HMR) have focused on pelvis-centered recovery, overlooking metric 3D localization and detection accuracy in the camera coordinate system - two key factors for real-world applications…
Human Mesh RecoveryScene UnderstandingBeyond Sin-Squared Error: Linear-Time Entrywise Uncertainty Quantification for Streaming PCA
We propose a novel statistical inference framework for streaming principal component analysis (PCA) using Oja's algorithm, enabling the construction of confidence intervals for individual entries of the estimated eigenve…
Uncertainty QuantificationKeep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language Models
Perspective-aware spatial reasoning involves understanding spatial relationships from specific viewpoints-either egocentric (observer-centered) or allocentric (object-centered). While vision-language models (VLMs) perfor…
Spatial ReasoningOpen-ended Hierarchical Streaming Video Understanding with Vision Language Models
We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datasets with hierarchical and fine-grained t…
Temporal Action LocalizationAction Classification