| Rank | Model |
Top-1 Accuracy | Top-5 Accuracy | Parameters | GFLOPs |
Extra Training Data |
Paper | Code | Year |
| 1 |
Decoupled Object-Centric Video Understan
자동 추출 |
86.79 | – | – | – |
|
Decoupled Object-Centric Video Understanding for Generating Robotic Manipulation Commands
|
|
2026 |
| 2 |
MVD (Kinetics400 pretrain, ViT-H, 16 frame) |
77.3 | 95.7 | 633 | 1192x6 |
✓ |
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning
|
ruiwang2021/mvd · 2023-MindSpore-4/Code-5 · Mind23-2/MindCode-3
· +1 |
2022 |
| 3 |
DejaVid |
77.2 | 96.3 | – | – |
✓ |
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification
|
darrylho/dejavid |
2025 |
| 3 |
InternVideo |
77.2 | – | – | – |
✓ |
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
|
opengvlab/internvideo · yingsen1/unimd |
2022 |
| 5 |
InternVideo2-1B |
77.1 | – | – | – |
✓ |
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
|
opengvlab/internvideo · opengvlab/internvideo2 |
2024 |
| 6 |
VideoMAE V2-g |
77.0 | 95.9 | 1013 | 2544x6 |
✓ |
VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
|
OpenGVLab/VideoMAEv2 |
2023 |
| 7 |
MVD (Kinetics400 pretrain, ViT-L, 16 frame) |
76.7 | 95.5 | 305 | 597x6 |
✓ |
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning
|
ruiwang2021/mvd · 2023-MindSpore-4/Code-5 · Mind23-2/MindCode-3
· +1 |
2022 |
| 8 |
Hiera-L (no extra data) |
76.5 | – | – | – |
|
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
|
huggingface/pytorch-image-models · facebookresearch/hiera · leondgarse/keras_cv_attention_models
· +1 |
2023 |
| 9 |
TubeViT-L |
76.1 | 95.2 | – | – |
|
Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning
|
daniel-code/TubeViT |
2022 |
| 10 |
VideoMAE (no extra data, ViT-L, 32x2) |
75.4 | 95.2 | 305 | 1436x3 |
|
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
|
huggingface/transformers · MCG-NJU/VideoMAE · MCG-NJU/VideoMAE-Action-Detection
· +6 |
2022 |
| 11 |
Side4Video (EVA ViT-E/14) |
75.2 | 94.0 | – | – |
|
Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning
|
whwu95/ATM · HJYao00/Side4Video |
2023 |
| 12 |
MaskFeat (Kinetics600 pretrain, MViT-L) |
75.0 | 95.0 | 218 | 2828*3 |
✓ |
Masked Feature Prediction for Self-Supervised Visual Pre-Training
|
facebookresearch/SlowFast · open-mmlab/mmselfsup · Westlake-AI/openmixup
· +3 |
2021 |
| 13 |
MAR (50% mask, ViT-L, 16x4) |
74.7 | 94.9 | 311 | 276x6 |
|
MAR: Masked Autoencoders for Efficient Action Recognition
|
alibaba-mmai-research/masked-action-recognition |
2022 |
| 14 |
ATM |
74.6 | 94.4 | – | – |
|
What Can Simple Arithmetic Operations Do for Temporal Modeling?
|
whwu95/ATM · HJYao00/Side4Video |
2023 |
| 15 |
MAWS (ViT-L) |
74.4 | – | – | – |
✓ |
The effectiveness of MAE pre-pretraining for billion-scale pretraining
|
facebookresearch/maws |
2023 |
| 16 |
VideoMAE (no extra data, ViT-L, 16frame) |
74.3 | 94.6 | 305 | 597x6 |
|
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
|
huggingface/transformers · MCG-NJU/VideoMAE · MCG-NJU/VideoMAE-Action-Detection
· +6 |
2022 |
| 17 |
MAR (75% mask, ViT-L, 16x4) |
73.8 | 94.4 | 311 | 131x6 |
|
MAR: Masked Autoencoders for Efficient Action Recognition
|
alibaba-mmai-research/masked-action-recognition |
2022 |
| 18 |
MVD (Kinetics400 pretrain, ViT-B, 16 frame) |
73.7 | 94.0 | 87 | 180x6 |
✓ |
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning
|
ruiwang2021/mvd · 2023-MindSpore-4/Code-5 · Mind23-2/MindCode-3
· +1 |
2022 |
| 18 |
ViC-MAE (ViT-L) |
73.7 | – | – | – |
|
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
|
jeffhernandez1995/vic-mae · MindCode-4/code-5 |
2023 |
| 20 |
TAdaFormer-L/14 |
73.6 | – | – | – |
✓ |
Temporally-Adaptive Models for Efficient Video Understanding
|
alibaba-mmai-research/TAdaConv |
2023 |