Relation Modeling in Spatio-Temporal Action Localization
This paper presents our solution to the AVA-Kinetics Crossover Challenge of ActivityNet workshop at CVPR 2021. Our solution utilizes multiple types of relation modeling methods for spatio-temporal action detection and adopts a training strategy to integrate multiple relation modeling in end-to-end training over the two large-scale video datasets. Learning with memory bank and finetuning for long-tailed distribution are also investigated to further improve the performance. In this paper, we detail the implementations of our solution and provide experiments results and corresponding discussions. We finally achieve 40.67 mAP on the test set of AVA-Kinetics.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionAction LocalizationRelationSpatio-Temporal Action LocalizationTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Modeling Spatio-Temporal Human Track Structure for Action Localization
This paper addresses spatio-temporal localization of human actions in video. In order to localize actions in time, we propose a recurrent localization network (RecLNet) designed to model the temporal structure of actions…
Action LocalizationHuman DetectionOptical Flow EstimationSpatio-Temporal Action Localization+2Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization
Localizing persons and recognizing their actions from videos is a challenging task towards high-level video understanding. Recent advances have been achieved by modeling direct pairwise relations between entities. In thi…
Action DetectionAction LocalizationAction RecognitionRelation+4CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Most current pipelines for spatio-temporal action localization connect frame-wise or clip-wise detection results to generate action proposals, where only local information is exploited and the efficiency is hindered by d…
Action DetectionAction LocalizationSpatio-Temporal Action LocalizationTemporal Action LocalizationMM-SEAL: A Large-scale Video Dataset of Multi-person Multi-grained Spatio-temporally Action Localization
In this paper, we introduce a novel large-scale video dataset dubbed MM-SEAL for multi-person multi-grained spatio-temporal action localization among human daily life. We are the first to propose a new benchmark for mult…
Action LocalizationAction RecognitionSpatio-Temporal Action LocalizationTemporal Action Localization+1TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection
Spatio-temporal action detection in videos requires jointly localizing actors in space and identifying action boundaries over time. A common challenge is constructing temporally stable action tubes, as frame-level detect…
Action Detection