Multi-region two-stream R-CNN for action detection
We propose a multi-region two-stream R-CNN model for action detection in realistic videos. We start from frame-level action detection based on faster R-CNN [1], and make three contributions: (1) we show that a motion region proposal network generates high-quality proposals , which are complementary to those of an appearance region proposal network; (2) we show that stacking optical flow over several frames significantly improves frame-level action detection; and (3) we embed a multi-region scheme in the faster R-CNN model, which adds complementary information on body parts. We then link frame-level detections with the Viterbi algorithm, and temporally localize an action with the maximum subarray method. Experimental results on the UCF-Sports, J-HMDB and UCF101 action detection datasets show that our approach outperforms the state of the art with a significant margin in both frame-mAP and video-mAP
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionAction RecognitionRegion ProposalSkeleton Based Action RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Improving Action Localization by Progressive Cross-stream Cooperation
Spatio-temporal action localization consists of three levels of tasks: spatial localization, action classification, and temporal segmentation. In this work, we propose a new Progressive Cross-stream Cooperation (PCSC) fr…
Action ClassificationAction DetectionAction LocalizationSpatio-Temporal Action Localization+1Assisted Refinement Network Based on Channel Information Interaction for Camouflaged and Salient Object Detection
Camouflaged Object Detection (COD) stands as a significant challenge in computer vision, dedicated to identifying and segmenting objects visually highly integrated with their backgrounds. Current mainstream methods have …
Salient Object DetectionObject LocalizationPolyp SegmentationSpectral-Spatial Synergistic Guided Network for Hyperspectral Salient Object Detection
Hyperspectral salient object detection aims to identify visually salient regions from hyperspectral images. Existing methods often fail because they fundamentally misunderstand the data, confusing incidental spectral var…
Computational EfficiencySalient Object DetectionLabel-Efficient Object Detection via Region Proposal Network Pre-Training
Self-supervised pre-training, based on the pretext task of instance discrimination, has fueled the recent advance in label-efficient object detection. However, existing studies focus on pre-training only a feature extrac…
Instance SegmentationObjectobject-detectionObject Detection+2Combining digital data streams and epidemic networks for real time outbreak detection
Responding to disease outbreaks requires close surveillance of their trajectories, but outbreak detection is hindered by the high noise in epidemic time series. Aggregating information across data sources has shown great…
Interpretable Machine Learning