paper-with-me

Action Recognition 벤치마크

Action Recognition on Something-Something V2

248개 결과 · ⬇ CSV · JSON

Top-1 Accuracy

1 22.45 43.9 65.34 86.79 2017-06 2026-09 model3D_1 with left-right augmentation and fps jitter — 51.33 (2017-06-13) model3D_1 with left-right augmentation and fps jitter — 51.33 (2017-06-13) TSM (RGB + Flow) — 66.6 (2018-11-20) TSM (RGB + Flow) — 66.6 (2018-11-20) SlowFast — 61.7 (2018-12-10) SlowFast — 61.7 (2018-12-10) Prob-Distill — 49.9 (2019-04-05) Prob-Distill — 49.9 (2019-04-05) TAM (5-shot) — 52.3 (2019-06-27) TAM (5-shot) — 52.3 (2019-06-27) TRG (ResNet-50) — 62.2 (2019-08-27) TRG (Inception-V3) — 61.3 (2019-08-27) CCS + two-stream + TRN — 61.2 (2019-08-27) TRG (ResNet-50) — 62.2 (2019-08-27) TRG (Inception-V3) — 61.3 (2019-08-27) CCS + two-stream + TRN — 61.2 (2019-08-27) STM + TRNMultiscale — 47.73 (2019-09-11) STM + TRNMultiscale — 47.73 (2019-09-11) bLVNet — 65.2 (2019-12-02) Multigrid — 61.7 (2019-12-02) bLVNet — 65.2 (2019-12-02) Multigrid — 61.7 (2019-12-02) TSM+W3 (16 frames, RGB ResNet-50) — 66.5 (2020-04-02) TSM+W3 (16 frames, RGB ResNet-50) — 66.5 (2020-04-02) TPN (TSM-50) — 62.0 (2020-04-07) TPN (TSM-50) — 62.0 (2020-04-07) MSNet-R50En (8+16 ensemble, ImageNet pretrained) — 66.6 (2020-07-20) MSNet-R50 (16 frames, ImageNet pretrained) — 64.7 (2020-07-20) MSNet-R50 (8 frames, ImageNet pretrained) — 63.0 (2020-07-20) MSNet-R50En (8+16 ensemble, ImageNet pretrained) — 66.6 (2020-07-20) MSNet-R50 (16 frames, ImageNet pretrained) — 64.7 (2020-07-20) MSNet-R50 (8 frames, ImageNet pretrained) — 63.0 (2020-07-20) PAN ResNet101 (RGB only, no Flow) — 66.5 (2020-08-08) PAN ResNet101 (RGB only, no Flow) — 66.5 (2020-08-08) MML (ensemble) — 69.02 (2020-11-04) MML (single) — 66.83 (2020-11-04) MML (ensemble) — 69.02 (2020-11-04) MML (single) — 66.83 (2020-11-04) VoV3D-L (32frames, Kinetics pretrained, single) — 67.35 (2020-12-01) VoV3D-L (32frames, from scratch, single) — 65.8 (2020-12-01) VoV3D-M (32frames, Kinetics pretrained, single) — 65.24 (2020-12-01) VoV3D-M (32frames, from scratch, single) — 64.2 (2020-12-01) VoV3D-L (16frames, from scratch, single) — 64.1 (2020-12-01) VoV3D-M (16frames, from scratch, single) — 63.2 (2020-12-01) VoV3D-L (32frames, Kinetics pretrained, single) — 67.35 (2020-12-01) VoV3D-L (32frames, from scratch, single) — 65.8 (2020-12-01) VoV3D-M (32frames, Kinetics pretrained, single) — 65.24 (2020-12-01) VoV3D-M (32frames, from scratch, single) — 64.2 (2020-12-01) VoV3D-L (16frames, from scratch, single) — 64.1 (2020-12-01) VoV3D-M (16frames, from scratch, single) — 63.2 (2020-12-01) MVFNet-ResNet50 (center crop, 8+16 ensemble, ImageNet pretrained, RGB only) — 66.3 (2020-12-13) MVFNet-ResNet50 (center crop, 8+16 ensemble, ImageNet pretrained, RGB only) — 66.3 (2020-12-13) TDN ResNet101 (one clip, three crop, 8+16 ensemble, ImageNet pretrained, RGB only) — 69.6 (2020-12-18) TDN ResNet101 (one clip, center crop, 8+16 ensemble, ImageNet pretrained, RGB only) — 68.2 (2020-12-18) TDN ResNet101 (one clip, three crop, 8+16 ensemble, ImageNet pretrained, RGB only) — 69.6 (2020-12-18) TDN ResNet101 (one clip, center crop, 8+16 ensemble, ImageNet pretrained, RGB only) — 68.2 (2020-12-18) TimeSformer-HR — 62.5 (2021-02-09) TimeSformer-L — 62.3 (2021-02-09) TimeSformer — 59.5 (2021-02-09) TimeSformer-HR — 62.5 (2021-02-09) TimeSformer-L — 62.3 (2021-02-09) TimeSformer — 59.5 (2021-02-09) SELFYNet-TSM-R50En (8+16 frames, ImageNet pretrained, 2 clips) — 67.7 (2021-02-14) SELFYNet-TSM-R50En (8+16 frames, ImageNet pretrained, a single clip) — 67.4 (2021-02-14) SELFYNet-TSM-R50 (16 frames, ImageNet pretrained) — 65.7 (2021-02-14) SELFYNet-TSM-R50En (8+16 frames, ImageNet pretrained, 2 clips) — 67.7 (2021-02-14) SELFYNet-TSM-R50En (8+16 frames, ImageNet pretrained, a single clip) — 67.4 (2021-02-14) SELFYNet-TSM-R50 (16 frames, ImageNet pretrained) — 65.7 (2021-02-14) MoViNet-A2 — 63.5 (2021-03-21) MoViNet-A1 — 62.7 (2021-03-21) MoViNet-A0 — 61.3 (2021-03-21) MoViNet-A2 — 63.5 (2021-03-21) MoViNet-A1 — 62.7 (2021-03-21) MoViNet-A0 — 61.3 (2021-03-21) ViViT-L/16x2 Fact. encoder — 65.4 (2021-03-29) ViViT-L/16x2 Fact. encoder — 65.4 (2021-03-29) MViT-B-24, 32x3 — 68.7 (2021-04-22) MViT-B, 32x3(Kinetics600 pretrain) — 67.8 (2021-04-22) MViT-B, 16x4 — 66.2 (2021-04-22) MViT-B-24, 32x3 — 68.7 (2021-04-22) MViT-B, 32x3(Kinetics600 pretrain) — 67.8 (2021-04-22) MViT-B, 16x4 — 66.2 (2021-04-22) VidTr-L — 60.2 (2021-04-23) VidTr-L — 60.2 (2021-04-23) CT-Net Ensemble (R50, 8+12+16+24) — 67.8 (2021-06-03) CT-Net Ensemble (R50, 8+12+16+24) — 67.8 (2021-06-03) Mformer-L — 68.1 (2021-06-09) Mformer-HR — 67.1 (2021-06-09) Mformer — 66.5 (2021-06-09) Mformer-L — 68.1 (2021-06-09) Mformer-HR — 67.1 (2021-06-09) Mformer — 66.5 (2021-06-09) X-Vit (x16) — 67.2 (2021-06-10) X-Vit (x16) — 67.2 (2021-06-10) VIMPAC — 68.1 (2021-06-21) VIMPAC — 68.1 (2021-06-21) Swin-B (IN-21K + Kinetics400 pretrain) — 69.6 (2021-06-24) Swin-B (IN-21K + Kinetics400 pretrain) — 69.6 (2021-06-24) UniFormer-B (IN-1K + Kinetics400 pretrain) — 71.2 (2021-09-29) UniFormer-S (IN-1K + Kinetics600 pretrain) — 69.4 (2021-09-29) UniFormer-B (IN-1K + Kinetics400 pretrain) — 71.2 (2021-09-29) UniFormer-S (IN-1K + Kinetics600 pretrain) — 69.4 (2021-09-29) TAda2D-En (ResNet-50, 8+16 frames) — 67.2 (2021-10-12) TAdaConvNeXt-T — 67.1 (2021-10-12) TAda2D (ResNet-50, 16 frames) — 65.6 (2021-10-12) TAda2D (ResNet-50, 8 frames) — 64.0 (2021-10-12) TAda2D-En (ResNet-50, 8+16 frames) — 67.2 (2021-10-12) TAdaConvNeXt-T — 67.1 (2021-10-12) TAda2D (ResNet-50, 16 frames) — 65.6 (2021-10-12) TAda2D (ResNet-50, 8 frames) — 64.0 (2021-10-12) ORViT Mformer-L (ORViT blocks) — 69.5 (2021-10-13) ORViT Mformer (ORViT blocks) — 67.9 (2021-10-13) ORViT Mformer-L (ORViT blocks) — 69.5 (2021-10-13) ORViT Mformer (ORViT blocks) — 67.9 (2021-10-13) RSANet-R50 (8+16 frames, ImageNet pretrained, 2 clips — 67.7 (2021-11-02) RSANet-R50 (8+16 frames, ImageNet pretrained, a single clip) — 67.3 (2021-11-02) RSANet-R50 (16 frames, ImageNet pretrained, a single clip) — 66.0 (2021-11-02) RSANet-R50 (8 frames, ImageNet pretrained, a single clip) — 64.8 (2021-11-02) RSANet-R50 (8+16 frames, ImageNet pretrained, 2 clips — 67.7 (2021-11-02) RSANet-R50 (8+16 frames, ImageNet pretrained, a single clip) — 67.3 (2021-11-02) RSANet-R50 (16 frames, ImageNet pretrained, a single clip) — 66.0 (2021-11-02) RSANet-R50 (8 frames, ImageNet pretrained, a single clip) — 64.8 (2021-11-02) MorphMLP-B (IN-1K) — 70.1 (2021-11-24) MorphMLP-B (IN-1K) — 70.1 (2021-11-24) MViTv2-L (IN-21K + Kinetics400 pretrain) — 73.3 (2021-12-02) MViT-B (IN-21K + Kinetics400 pretrain) — 72.1 (2021-12-02) BEVT (IN-1K + Kinetics400 pretrain) — 71.4 (2021-12-02) SVT — 59.2 (2021-12-02) MViTv2-L (IN-21K + Kinetics400 pretrain) — 73.3 (2021-12-02) MViT-B (IN-21K + Kinetics400 pretrain) — 72.1 (2021-12-02) BEVT (IN-1K + Kinetics400 pretrain) — 71.4 (2021-12-02) SVT — 59.2 (2021-12-02) CoVeR(JFT-3B) — 70.9 (2021-12-14) CoVeR(JFT-300M) — 69.8 (2021-12-14) CoVeR(JFT-3B) — 70.9 (2021-12-14) CoVeR(JFT-300M) — 69.8 (2021-12-14) MaskFeat (Kinetics600 pretrain, MViT-L) — 75.0 (2021-12-16) MaskFeat (Kinetics600 pretrain, MViT-L) — 75.0 (2021-12-16) MTV-B — 68.5 (2022-01-12) MTV-B — 68.5 (2022-01-12) AK-Net — 64.3 (2022-01-17) AK-Net — 64.3 (2022-01-17) OMNIVORE (Swin-B, IN-21K+ Kinetics400 pretrain) — 71.4 (2022-01-20) OMNIVORE (Swin-B, IN-21K+ Kinetics400 pretrain) — 71.4 (2022-01-20) TCM (Ensemble) — 67.8 (2022-02-24) TCM (Ensemble) — 67.8 (2022-02-24) GC-TDN Ensemble (R50,8+16) — 67.8 (2022-03-18) GC-TDN Ensemble (R50,8+16) — 67.8 (2022-03-18) DirecFormer — 64.94 (2022-03-19) DirecFormer — 64.94 (2022-03-19) VideoMAE (no extra data, ViT-L, 32x2) — 75.4 (2022-03-23) VideoMAE (no extra data, ViT-L, 16frame) — 74.3 (2022-03-23) VideoMAE (no extra data, ViT-B, 16frame) — 70.8 (2022-03-23) VideoMAE (no extra data, ViT-L, 32x2) — 75.4 (2022-03-23) VideoMAE (no extra data, ViT-L, 16frame) — 74.3 (2022-03-23) VideoMAE (no extra data, ViT-B, 16frame) — 70.8 (2022-03-23) MLP-3D — 68.5 (2022-06-13) MLP-3D — 68.5 (2022-06-13) SIFA — 69.8 (2022-06-14) SIFA — 69.8 (2022-06-14) ST-Adapter (ViT-L, CLIP) — 72.3 (2022-06-27) ST-Adapter (ViT-L, CLIP) — 72.3 (2022-06-27) MoDS (8+16frames) — 67.1 (2022-07-15) MoDS (8+16frames) — 67.1 (2022-07-15) MAR (50% mask, ViT-L, 16x4) — 74.7 (2022-07-24) MAR (75% mask, ViT-L, 16x4) — 73.8 (2022-07-24) MAR (50% mask, ViT-B, 16x4) — 71.0 (2022-07-24) MAR (75% mask, ViT-B, 16x4) — 69.5 (2022-07-24) MAR (50% mask, ViT-L, 16x4) — 74.7 (2022-07-24) MAR (75% mask, ViT-L, 16x4) — 73.8 (2022-07-24) MAR (50% mask, ViT-B, 16x4) — 71.0 (2022-07-24) MAR (75% mask, ViT-B, 16x4) — 69.5 (2022-07-24) TPS — 69.8 (2022-07-27) TPS — 69.8 (2022-07-27) STPG (8+16frames) — 67.0 (2022-08-09) STPG (8+16frames) — 67.0 (2022-08-09) OmniVL — 62.5 (2022-09-15) OmniVL — 62.5 (2022-09-15) UniFormerV2-L — 73.0 (2022-09-22) UniFormerV2-L — 73.0 (2022-09-22) GTDNet — 67.6 (2022-11-23) GTDNet — 67.6 (2022-11-23) InternVideo — 77.2 (2022-12-06) TubeViT-L — 76.1 (2022-12-06) InternVideo — 77.2 (2022-12-06) TubeViT-L — 76.1 (2022-12-06) MVD (Kinetics400 pretrain, ViT-H, 16 frame) — 77.3 (2022-12-08) MVD (Kinetics400 pretrain, ViT-L, 16 frame) — 76.7 (2022-12-08) MVD (Kinetics400 pretrain, ViT-B, 16 frame) — 73.7 (2022-12-08) MVD (Kinetics400 pretrain, ViT-S, 16 frame) — 70.9 (2022-12-08) MVD (Kinetics400 pretrain, ViT-H, 16 frame) — 77.3 (2022-12-08) MVD (Kinetics400 pretrain, ViT-L, 16 frame) — 76.7 (2022-12-08) MVD (Kinetics400 pretrain, ViT-B, 16 frame) — 73.7 (2022-12-08) MVD (Kinetics400 pretrain, ViT-S, 16 frame) — 70.9 (2022-12-08) MSMA (8+16frames) — 68.2 (2023-02-19) MSMA (8+16frames) — 68.2 (2023-02-19) E3D-L — 65.7 (2023-03-05) E3D-L — 65.7 (2023-03-05) ViC-MAE (ViT-L) — 73.7 (2023-03-21) ViC-MAE (ViT-L) — 73.7 (2023-03-21) MAWS (ViT-L) — 74.4 (2023-03-23) MAWS (ViT-L) — 74.4 (2023-03-23) VideoMAE V2-g — 77.0 (2023-03-29) VideoMAE V2-g — 77.0 (2023-03-29) ILA (ViT-L/14) — 70.2 (2023-04-20) ILA (ViT-B/16) — 66.8 (2023-04-20) ILA (ViT-L/14) — 70.2 (2023-04-20) ILA (ViT-B/16) — 66.8 (2023-04-20) PLAR — 67.3 (2023-05-21) PLAR — 67.3 (2023-05-21) Hiera-L (no extra data) — 76.5 (2023-06-01) Hiera-L (no extra data) — 76.5 (2023-06-01) ATM — 74.6 (2023-07-18) ATM — 74.6 (2023-07-18) TAdaFormer-L/14 — 73.6 (2023-08-10) TAdaConvNeXtV2-B — 71.1 (2023-08-10) TAdaFormer-L/14 — 73.6 (2023-08-10) TAdaConvNeXtV2-B — 71.1 (2023-08-10) ZeroI2V ViT-L/14 — 72.2 (2023-10-02) ZeroI2V ViT-L/14 — 72.2 (2023-10-02) AMD(ViT-B/16) — 73.3 (2023-11-06) AMD(ViT-S/16) — 70.2 (2023-11-06) AMD(ViT-B/16) — 73.3 (2023-11-06) AMD(ViT-S/16) — 70.2 (2023-11-06) Side4Video (EVA ViT-E/14) — 75.2 (2023-11-27) Side4Video (EVA ViT-E/14) — 75.2 (2023-11-27) CAST(ViT-B/16) — 71.6 (2023-11-30) CAST(ViT-B/16) — 71.6 (2023-11-30) InternVideo2-1B — 77.1 (2024-03-22) InternVideo2-6B — 1.0 (2024-03-22) InternVideo2-1B — 77.1 (2024-03-22) InternVideo2-6B — 1.0 (2024-03-22) StructVit-B-4-1 — 71.5 (2024-04-05) StructVit-B-4-1 — 71.5 (2024-04-05) TDS-CLIP-ViT-L/14(8frames) — 73.4 (2024-08-20) TDS-CLIP-ViT-L/14(8frames) — 73.4 (2024-08-20) DejaVid — 77.2 (2025-01-01) DejaVid — 77.2 (2025-01-01) TrAction — 45.0 (2026-06-02) Decoupled Object-Centric Video Understan — 86.79 (2026-06-15) model3D_1 with left-right augmentation and fps jitter — 51.33 (2017-06-13) TSM (RGB + Flow) — 66.6 (2018-11-20) MML (ensemble) — 69.02 (2020-11-04) TDN ResNet101 (one clip, three crop, 8+16 ensemble, ImageNet pretrained, RGB only) — 69.6 (2020-12-18) UniFormer-B (IN-1K + Kinetics400 pretrain) — 71.2 (2021-09-29) MViTv2-L (IN-21K + Kinetics400 pretrain) — 73.3 (2021-12-02) MaskFeat (Kinetics600 pretrain, MViT-L) — 75.0 (2021-12-16) VideoMAE (no extra data, ViT-L, 32x2) — 75.4 (2022-03-23) InternVideo — 77.2 (2022-12-06) MVD (Kinetics400 pretrain, ViT-H, 16 frame) — 77.3 (2022-12-08) Decoupled Object-Centric Video Understan — 86.79 (2026-06-15)
RankModel Top-1 AccuracyTop-5 AccuracyParametersGFLOPs Extra Training Data PaperCodeYear
1 Decoupled Object-Centric Video Understan 자동 추출 86.79 Decoupled Object-Centric Video Understanding for Generating Robotic Manipulation Commands 2026
2 MVD (Kinetics400 pretrain, ViT-H, 16 frame) 77.395.76331192x6 Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning ruiwang2021/mvd · 2023-MindSpore-4/Code-5 · Mind23-2/MindCode-3 · +1 2022
3 DejaVid 77.296.3 DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification darrylho/dejavid 2025
3 InternVideo 77.2 InternVideo: General Video Foundation Models via Generative and Discriminative Learning opengvlab/internvideo · yingsen1/unimd 2022
5 InternVideo2-1B 77.1 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding opengvlab/internvideo · opengvlab/internvideo2 2024
6 VideoMAE V2-g 77.095.910132544x6 VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking OpenGVLab/VideoMAEv2 2023
7 MVD (Kinetics400 pretrain, ViT-L, 16 frame) 76.795.5305597x6 Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning ruiwang2021/mvd · 2023-MindSpore-4/Code-5 · Mind23-2/MindCode-3 · +1 2022
8 Hiera-L (no extra data) 76.5 Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles huggingface/pytorch-image-models · facebookresearch/hiera · leondgarse/keras_cv_attention_models · +1 2023
9 TubeViT-L 76.195.2 Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning daniel-code/TubeViT 2022
10 VideoMAE (no extra data, ViT-L, 32x2) 75.495.23051436x3 VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training huggingface/transformers · MCG-NJU/VideoMAE · MCG-NJU/VideoMAE-Action-Detection · +6 2022
11 Side4Video (EVA ViT-E/14) 75.294.0 Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning whwu95/ATM · HJYao00/Side4Video 2023
12 MaskFeat (Kinetics600 pretrain, MViT-L) 75.095.02182828*3 Masked Feature Prediction for Self-Supervised Visual Pre-Training facebookresearch/SlowFast · open-mmlab/mmselfsup · Westlake-AI/openmixup · +3 2021
13 MAR (50% mask, ViT-L, 16x4) 74.794.9311276x6 MAR: Masked Autoencoders for Efficient Action Recognition alibaba-mmai-research/masked-action-recognition 2022
14 ATM 74.694.4 What Can Simple Arithmetic Operations Do for Temporal Modeling? whwu95/ATM · HJYao00/Side4Video 2023
15 MAWS (ViT-L) 74.4 The effectiveness of MAE pre-pretraining for billion-scale pretraining facebookresearch/maws 2023
16 VideoMAE (no extra data, ViT-L, 16frame) 74.394.6305597x6 VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training huggingface/transformers · MCG-NJU/VideoMAE · MCG-NJU/VideoMAE-Action-Detection · +6 2022
17 MAR (75% mask, ViT-L, 16x4) 73.894.4311131x6 MAR: Masked Autoencoders for Efficient Action Recognition alibaba-mmai-research/masked-action-recognition 2022
18 MVD (Kinetics400 pretrain, ViT-B, 16 frame) 73.794.087180x6 Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning ruiwang2021/mvd · 2023-MindSpore-4/Code-5 · Mind23-2/MindCode-3 · +1 2022
18 ViC-MAE (ViT-L) 73.7 ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders jeffhernandez1995/vic-mae · MindCode-4/code-5 2023
20 TAdaFormer-L/14 73.6 Temporally-Adaptive Models for Efficient Video Understanding alibaba-mmai-research/TAdaConv 2023
1–20 / 248 다음 → 페이지당 10 20 50 100