Point-Level Temporal Action Localization: Bridging Fully-supervised Proposals to Weakly-supervised Losses
Point-Level temporal action localization (PTAL) aims to localize actions in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the frame-level prediction paradigm to learn from the sparse single-frame labels. However, such a framework inevitably suffers from a large solution space. This paper attempts to explore the proposal-based prediction paradigm for point-level annotations, which has the advantage of more constrained solution space and consistent predictions among neighboring frames. The point-level annotations are first used as the keypoint supervision to train a keypoint detector. At the location prediction stage, a simple but effective mapper module, which enables back-propagation of training errors, is then introduced to bridge the fully-supervised framework with weak supervision. To our best of knowledge, this is the first work to leverage the fully-supervised paradigm for the point-level setting. Experiments on THUMOS14, BEOID, and GTEA verify the effectiveness of our proposed method both quantitatively and qualitatively, and demonstrate that our method outperforms state-of-the-art methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Action LocalizationPredictionTemporal Action LocalizationWeakly Supervised Action LocalizationSimilar Papers 제목 키워드 기반
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per action instance) to train a model to effectively locate action instances wit…
Weakly-supervised Temporal Action LocalizationMulti-Task LearningOnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
Temporal Action Localization (TAL) typically relies on segment annotations or offline access to full videos, limiting scalability and online use. We introduce Point-Supervised Online TAL (POTAL), which localizes actions …
Temporal Action LocalizationWeakly Supervised Temporal Action Localization via Representative Snippet Knowledge Propagation
Weakly supervised temporal action localization aims to localize temporal boundaries of actions and simultaneously identify their categories with only video-level category labels. Many existing methods seek to generate ps…
Action LocalizationPseudo LabelTemporal Action LocalizationWeakly-supervised Temporal Action LocalizationBoosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal Transport
Point-supervised Temporal Action Localization poses significant challenges due to the difficulty of identifying complete actions with a single-point annotation per action. Existing methods typically employ Multiple …
Action LocalizationMultiple Instance LearningTemporal Action LocalizationSub-action Prototype Learning for Point-level Weakly-supervised Temporal Action Localization
Point-level weakly-supervised temporal action localization (PWTAL) aims to localize actions with only a single timestamp annotation for each action instance. Existing methods tend to mine dense pseudo labels to alleviate…
Action LocalizationPseudo LabelTemporal Action LocalizationWeakly-supervised Temporal Action Localization