Weakly-supervised Action Localization with Background Modeling
We describe a latent approach that learns to detect actions in long sequences given training videos with only whole-video class labels. Our approach makes use of two innovations to attention-modeling in weakly-supervised learning. First, and most notably, our framework uses an attention model to extract both foreground and background frames whose appearance is explicitly modeled. Most prior works ignore the background, but we show that modeling it allows our system to learn a richer notion of actions and their temporal extents. Second, we combine bottom-up, class-agnostic attention modules with top-down, class-specific activation maps, using the latter as form of self-supervision for the former. Doing so allows our model to learn a more accurate model of attention without explicit temporal supervision. These modifications lead to 10% AP@IoU=0.5 improvement over existing systems on THUMOS14. Our proposed weaklysupervised system outperforms recent state-of-the-arts by at least 4.3% AP@IoU=0.5. Finally, we demonstrate that weakly-supervised learning can be used to aggressively scale-up learning to in-the-wild, uncurated Instagram videos. The addition of these videos significantly improves localization performance of our weakly-supervised model
Code (0)
등록된 구현이 없습니다.
Tasks
Action LocalizationWeakly Supervised Action LocalizationWeakly-supervised LearningSimilar Papers 제목 키워드 기반
Weakly-Supervised Temporal Action Localization Through Local-Global Background Modeling
Weakly-Supervised Temporal Action Localization (WS-TAL) task aims to recognize and localize temporal starts and ends of action instances in an untrimmed video with only video-level label supervision. Due to lack of negat…
Action LocalizationTemporal Action LocalizationWeakly-supervised LearningWeakly-supervised Temporal Action LocalizationWeakly-supervised Temporal Action Localization by Uncertainty Modeling
Weakly-supervised temporal action localization aims to learn detecting temporal intervals of action classes with only video-level labels. To this end, it is crucial to separate frames of action classes from the backgroun…
Action ClassificationAction LocalizationMultiple Instance LearningOut-of-Distribution Detection+3Background-Click Supervision for Temporal Action Localization
Weakly supervised temporal action localization aims at learning the instance-level action pattern from the video-level labels, where a significant challenge is action-context confusion. To overcome this challenge, one re…
Action LocalizationPositionTemporal Action LocalizationWeakly-supervised Temporal Action LocalizationForcing the Whole Video as Background: An Adversarial Learning Strategy for Weakly Temporal Action Localization
With video-level labels, weakly supervised temporal action localization (WTAL) applies a localization-by-classification paradigm to detect and classify the action in untrimmed videos. Due to the characteristic of classif…
Action LocalizationClassificationTemporal Action LocalizationWeakly-supervised Temporal Action LocalizationACM-Net: Action Context Modeling Network for Weakly-Supervised Temporal Action Localization
Weakly-supervised temporal action localization aims to localize action instances temporal boundary and identify the corresponding action category with only video-level labels. Traditional methods mainly focus on foregrou…
Action LocalizationTemporal Action LocalizationWeakly Supervised Action LocalizationWeakly-supervised Temporal Action Localization