Boundary-sensitive Pre-training for Temporal Localization in Videos
Many video analysis tasks require temporal localization thus detection of content changes. However, most existing models developed for these tasks are pre-trained on general video action classification tasks. This is because large scale annotation of temporal boundaries in untrimmed videos is expensive. Therefore no suitable datasets exist for temporal boundary-sensitive pre-training. In this paper for the first time, we investigate model pre-training for temporal localization by introducing a novel boundary-sensitive pretext (BSP) task. Instead of relying on costly manual annotations of temporal boundaries, we propose to synthesize temporal boundaries in existing video action classification datasets. With the synthesized boundaries, BSP can be simply conducted via classifying the boundary types. This enables the learning of video representations that are much more transferable to downstream temporal localization tasks. Extensive experiments show that the proposed BSP is superior and complementary to the existing action classification based pre-training counterpart, and achieves new state-of-the-art performance on several temporal localization tasks.
Code (1)
Tasks
Action ClassificationClassificationGeneral ClassificationTemporal Action LocalizationTemporal LocalizationSimilar Papers 제목 키워드 기반
Faster Learning of Temporal Action Proposal via Sparse Multilevel Boundary Generator
Temporal action localization in videos presents significant challenges in the field of computer vision. While the boundary-sensitive method has been widely adopted, its limitations include incomplete use of intermediate …
Action LocalizationTemporal Action LocalizationExplainable Forensics of Manipulated Segments in Untrimmed Long Videos
The rapid advancement of AI-driven video generation has transformed content creation, while simultaneously increasing the risk of misinformation through localized manipulations in long-form videos. Existing video forensi…
Video GenerationPcmNet: Position-Sensitive Context Modeling Network for Temporal Action Localization
Temporal action localization is an important and challenging task that aims to locate temporal regions in real-world untrimmed videos where actions occur and recognize their classes. It is widely acknowledged that video …
Action LocalizationBoundary DetectionPositionTemporal Action Localization+2Anchor-free temporal action localization via Progressive Boundary-aware Boosting
Enormous untrimmed videos from the real world are difficult to analyze and manage. Temporal action localization algorithms can help us to locate and recognize human activity clips in untrimmed videos. Recently, anchor-fr…
Action LocalizationTemporal Action LocalizationTemporal Perceiving Video-Language Pre-training
Video-Language Pre-training models have recently significantly improved various multi-modal downstream tasks. Previous dominant works mainly adopt contrastive learning to achieve global feature alignment across modalitie…
Action LocalizationContrastive LearningMoment RetrievalQuestion Answering+6