Temporal Action Localization with Enhanced Instant Discriminability
Temporal action detection (TAD) aims to detect all action boundaries and their corresponding categories in an untrimmed video. The unclear boundaries of actions in videos often result in imprecise predictions of action boundaries by existing methods. To resolve this issue, we propose a one-stage framework named TriDet. First, we propose a Trident-head to model the action boundary via an estimated relative probability distribution around the boundary. Then, we analyze the rank-loss problem (i.e. instant discriminability deterioration) in transformer-based methods and propose an efficient scalable-granularity perception (SGP) layer to mitigate this issue. To further push the limit of instant discriminability in the video backbone, we leverage the strong representation capability of pretrained large models and investigate their performance on TAD. Last, considering the adequate spatial-temporal context for classification, we design a decoupled feature pyramid network with separate feature pyramids to incorporate rich spatial context from the large model for localization. Experimental results demonstrate the robustness of TriDet and its state-of-the-art performance on multiple TAD datasets, including hierarchical (multilabel) TAD datasets.
Code (3)
Tasks
Action DetectionAction LocalizationTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
Most research efforts in the multimedia forensics domain have focused on detecting forgery audio-visual content and reached sound achievements. However, these works only consider deepfake detection as a classification ta…
Anomaly DetectionContrastive LearningDeepFake DetectionFace Swapping+13C-Net: Category Count and Center Loss for Weakly-Supervised Action Localization
Temporal action localization is a challenging computer vision problem with numerous real-world applications. Most existing methods require laborious frame-level supervision to train action localization models. In this wo…
Action ClassificationAction LocalizationTemporal Action LocalizationWeakly Supervised Action Localization+1DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action Localization
Weakly-supervised temporal action localization (WTAL) is a practical yet challenging task. Due to large-scale datasets, most existing methods use a network pretrained in other datasets to extract features, which are not …
Action LocalizationTemporal Action LocalizationWeakly-supervised Temporal Action LocalizationJCDNet: Joint of Common and Definite phases Network for Weakly Supervised Temporal Action Localization
Weakly-supervised temporal action localization aims to localize action instances in untrimmed videos with only video-level supervision. We witness that different actions record common phases, e.g., the run-up in the High…
Action LocalizationMultiple Instance LearningTemporal Action LocalizationWeakly-supervised Learning+1Forcing the Whole Video as Background: An Adversarial Learning Strategy for Weakly Temporal Action Localization
With video-level labels, weakly supervised temporal action localization (WTAL) applies a localization-by-classification paradigm to detect and classify the action in untrimmed videos. Due to the characteristic of classif…
Action LocalizationClassificationTemporal Action LocalizationWeakly-supervised Temporal Action Localization