Spatio-Temporal Action Localization 벤치마크
Spatio-Temporal Action Localization on AVA-Kinetics
val mAP
- 2020-06-14 — ACAR (multi-scale, ensemble): val mAP 40.49
- 2021-06-15 — RM (multi-scale, ensemble): val mAP 40.52
- 2022-12-06 — InternVideo: val mAP 41.01
- 2023-03-29 — VideoMAE V2-g: val mAP 42.6
| Rank | Model | val mAP | test mAP | Extra Training Data | Paper | Code | Year |
|---|---|---|---|---|---|---|---|
| 1 | VideoMAE V2-g | 42.6 | – | ✓ | VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking | OpenGVLab/VideoMAEv2 | 2023 |
| 2 | STAR/L | 41.7 | – | ✓ | End-to-End Spatio-Temporal Action Localisation with Video Transformers | 2023 | |
| 3 | InternVideo | 41.01 | – | ✓ | InternVideo: General Video Foundation Models via Generative and Discriminative Learning | opengvlab/internvideo · yingsen1/unimd | 2022 |
| 4 | RM (multi-scale, ensemble) | 40.52 | – | ✓ | Relation Modeling in Spatio-Temporal Action Localization | 2021 | |
| 5 | ACAR (multi-scale, ensemble) | 40.49 | 39.62 | ✓ | Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization | towhee-io/towhee · Siyu-C/ACAR-Net · salmank255/ROADSlowFast | 2020 |
| 6 | RM (multi-scale, ir-CSN-152) | 37.95 | – | Relation Modeling in Spatio-Temporal Action Localization | 2021 | ||
| 7 | ACAR (multi-scale, R-101, 8 × 8) | 36.36 | – | Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization | towhee-io/towhee · Siyu-C/ACAR-Net · salmank255/ROADSlowFast | 2020 |