Two-Stream Consensus Network: Submission to HACS Challenge 2021 Weakly-Supervised Learning Track
This technical report presents our solution to the HACS Temporal Action Localization Challenge 2021, Weakly-Supervised Learning Track. The goal of weakly-supervised temporal action localization is to temporally locate and classify action of interest in untrimmed videos given only video-level labels. We adopt the two-stream consensus network (TSCN) as the main framework in this challenge. The TSCN consists of a two-stream base model training procedure and a pseudo ground truth learning procedure. The base model training encourages the model to predict reliable predictions based on single modality (i.e., RGB or optical flow), based on the fusion of which a pseudo ground truth is generated and in turn used as supervision to train the base models. On the HACS v1.1.1 dataset, without fine-tuning the feature-extraction I3D models, our method achieves 22.20% on the validation set and 21.68% on the testing set in terms of average mAP. Our solution ranked the 2nd in this challenge, and we hope our method can serve as a baseline for future academic research.
Code (0)
등록된 구현이 없습니다.
Tasks
Action LocalizationOptical Flow EstimationTemporal Action LocalizationWeakly-supervised LearningWeakly-supervised Temporal Action LocalizationSimilar Papers 제목 키워드 기반
Transferable Knowledge-Based Multi-Granularity Aggregation Network for Temporal Action Localization: Submission to ActivityNet Challenge 2021
This technical report presents an overview of our solution used in the submission to 2021 HACS Temporal Action Localization Challenge on both Supervised Learning Track and Weakly-Supervised Learning Track. Temporal Actio…
Action LocalizationTemporal Action LocalizationTransfer LearningWeakly-supervised Learning+1HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization
This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and…
Action ClassificationAction LocalizationAction RecognitionTemporal Action Localization+2Weakly-Supervised Temporal Action Localization Through Local-Global Background Modeling
Weakly-Supervised Temporal Action Localization (WS-TAL) task aims to recognize and localize temporal starts and ends of action instances in an untrimmed video with only video-level label supervision. Due to lack of negat…
Action LocalizationTemporal Action LocalizationWeakly-supervised LearningWeakly-supervised Temporal Action LocalizationTwo-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
Weakly-supervised Temporal Action Localization (W-TAL) aims to classify and localize all action instances in an untrimmed video under only video-level supervision. However, without frame-level annotations, it is challeng…
Action LocalizationTemporal Action LocalizationVocal Bursts Valence PredictionWeakly Supervised Action Localization+1STAT: Towards Generalizable Temporal Action Localization
Weakly-supervised temporal action localization (WTAL) aims to recognize and localize action instances with only video-level labels. Despite the significant progress, existing methods suffer from severe performance degrad…
Action LocalizationTemporal Action LocalizationWeakly-supervised Temporal Action Localization