paper-with-me

Action Recognition 벤치마크

Action Recognition on Diving-48

36개 결과 · ⬇ CSV · JSON

Accuracy

75 79.97 84.95 89.93 94.9 2018-12 2026-09 SlowFast — 77.6 (2018-12-10) SlowFast — 77.6 (2018-12-10) TimeSformer-L — 81.0 (2021-02-09) TimeSformer-HR — 78.0 (2021-02-09) TimeSformer — 75.0 (2021-02-09) TimeSformer-L — 81.0 (2021-02-09) TimeSformer-HR — 78.0 (2021-02-09) TimeSformer — 75.0 (2021-02-09) TQN — 81.8 (2021-04-19) TQN — 81.8 (2021-04-19) VIMPAC — 85.5 (2021-06-21) VIMPAC — 85.5 (2021-06-21) ORViT TimeSformer — 88.0 (2021-10-13) ORViT TimeSformer — 88.0 (2021-10-13) RSANet-R50 (16 frames, ImageNet pretrained, a single clip) — 84.2 (2021-11-02) RSANet-R50 (16 frames, ImageNet pretrained, a single clip) — 84.2 (2021-11-02) BEVT — 86.7 (2021-12-02) BEVT — 86.7 (2021-12-02) TFCNet — 88.3 (2022-03-11) TFCNet — 88.3 (2022-03-11) GC-TDN — 87.6 (2022-03-18) GC-TDN — 87.6 (2022-03-18) PSB — 86.0 (2022-07-27) PSB — 86.0 (2022-07-27) AIM (CLIP ViT-L/14, 32x224) — 90.6 (2023-02-06) AIM (CLIP ViT-L/14, 32x224) — 90.6 (2023-02-06) DUALPATH — 88.7 (2023-03-17) DUALPATH — 88.7 (2023-03-17) PMI Sampler — 81.3 (2023-04-14) PMI Sampler — 81.3 (2023-04-14) Video-FocalNet-B — 90.8 (2023-07-13) Video-FocalNet-B — 90.8 (2023-07-13) StructVit-B-4-1 — 88.3 (2024-04-05) StructVit-B-4-1 — 88.3 (2024-04-05) LVMAE — 94.9 (2024-11-20) LVMAE — 94.9 (2024-11-20) SlowFast — 77.6 (2018-12-10) TimeSformer-L — 81.0 (2021-02-09) TQN — 81.8 (2021-04-19) VIMPAC — 85.5 (2021-06-21) ORViT TimeSformer — 88.0 (2021-10-13) TFCNet — 88.3 (2022-03-11) AIM (CLIP ViT-L/14, 32x224) — 90.6 (2023-02-06) Video-FocalNet-B — 90.8 (2023-07-13) LVMAE — 94.9 (2024-11-20)
RankModel Accuracy Extra Training Data PaperCodeYear
1 LVMAE 94.9 Extending Video Masked Autoencoders to 128 frames 2024
2 Video-FocalNet-B 90.8 Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition talalwasim/video-focalnets · innat/Video-FocalNets · hayatkhan8660-maker/DVFL-Net 2023
3 AIM (CLIP ViT-L/14, 32x224) 90.6 AIM: Adapting Image Models for Efficient Video Action Recognition taoyang1122/adapt-image-models 2023
4 DUALPATH 88.7 Dual-path Adaptation from Image to Video Transformers park-jungin/dualpath 2023
5 TFCNet 88.3 TFCNet: Temporal Fully Connected Networks for Static Unbiased Temporal Reasoning 2022
5 StructVit-B-4-1 88.3 Learning Correlation Structures for Vision Transformers 2024
7 ORViT TimeSformer 88.0 Object-Region Video Transformers eladb3/orvit 2021
8 GC-TDN 87.6 Group Contextualization for Video Recognition haoyanbin918/group-contextualization 2022
9 BEVT 86.7 BEVT: BERT Pretraining of Video Transformers xyzforever/bevt 2021
10 PSB 86 Spatiotemporal Self-attention Modeling with Temporal Patch Shift for Action Recognition martinxm/tps 2022
11 VIMPAC 85.5 VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning airsplay/vimpac 2021
12 RSANet-R50 (16 frames, ImageNet pretrained, a single clip) 84.2 Relational Self-Attention: What's Missing in Attention for Video Understanding KimManjin/RSA 2021
13 TQN 81.8 Temporal Query Networks for Fine-grained Video Understanding 2021
14 PMI Sampler 81.3 PMI Sampler: Patch Similarity Guided Frame Selection for Aerial Action Recognition ricky-xian/pmi-sampler 2023
15 TimeSformer-L 81 Is Space-Time Attention All You Need for Video Understanding? open-mmlab/mmaction2 · towhee-io/towhee · facebookresearch/TimeSformer · +13 2021
16 TimeSformer-HR 78 Is Space-Time Attention All You Need for Video Understanding? open-mmlab/mmaction2 · towhee-io/towhee · facebookresearch/TimeSformer · +13 2021
17 SlowFast 77.6 SlowFast Networks for Video Recognition facebookresearch/SlowFast · open-mmlab/mmaction2 · facebookresearch/pytorchvideo · +12 2018
18 TimeSformer 75 Is Space-Time Attention All You Need for Video Understanding? open-mmlab/mmaction2 · towhee-io/towhee · facebookresearch/TimeSformer · +13 2021
19 LVMAE 94.9 Extending Video Masked Autoencoders to 128 frames 2024
20 Video-FocalNet-B 90.8 Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition talalwasim/video-focalnets · innat/Video-FocalNets · hayatkhan8660-maker/DVFL-Net 2023
21 AIM (CLIP ViT-L/14, 32x224) 90.6 AIM: Adapting Image Models for Efficient Video Action Recognition taoyang1122/adapt-image-models 2023
22 DUALPATH 88.7 Dual-path Adaptation from Image to Video Transformers park-jungin/dualpath 2023
23 TFCNet 88.3 TFCNet: Temporal Fully Connected Networks for Static Unbiased Temporal Reasoning 2022
23 StructVit-B-4-1 88.3 Learning Correlation Structures for Vision Transformers 2024
25 ORViT TimeSformer 88.0 Object-Region Video Transformers eladb3/orvit 2021
26 GC-TDN 87.6 Group Contextualization for Video Recognition haoyanbin918/group-contextualization 2022
27 BEVT 86.7 BEVT: BERT Pretraining of Video Transformers xyzforever/bevt 2021
28 PSB 86 Spatiotemporal Self-attention Modeling with Temporal Patch Shift for Action Recognition martinxm/tps 2022
29 VIMPAC 85.5 VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning airsplay/vimpac 2021
30 RSANet-R50 (16 frames, ImageNet pretrained, a single clip) 84.2 Relational Self-Attention: What's Missing in Attention for Video Understanding KimManjin/RSA 2021
31 TQN 81.8 Temporal Query Networks for Fine-grained Video Understanding 2021
32 PMI Sampler 81.3 PMI Sampler: Patch Similarity Guided Frame Selection for Aerial Action Recognition ricky-xian/pmi-sampler 2023
33 TimeSformer-L 81 Is Space-Time Attention All You Need for Video Understanding? open-mmlab/mmaction2 · towhee-io/towhee · facebookresearch/TimeSformer · +13 2021
34 TimeSformer-HR 78 Is Space-Time Attention All You Need for Video Understanding? open-mmlab/mmaction2 · towhee-io/towhee · facebookresearch/TimeSformer · +13 2021
35 SlowFast 77.6 SlowFast Networks for Video Recognition facebookresearch/SlowFast · open-mmlab/mmaction2 · facebookresearch/pytorchvideo · +12 2018
36 TimeSformer 75 Is Space-Time Attention All You Need for Video Understanding? open-mmlab/mmaction2 · towhee-io/towhee · facebookresearch/TimeSformer · +13 2021
1–36 / 36 페이지당 10 20 50 100