paper-with-me

Image Classification 벤치마크

Image Classification on ObjectNet

106개 결과 · ⬇ CSV · JSON

Top-1 Accuracy

4.92 24.37 43.81 63.26 82.7 2019-10 2026-09 ResNet-50 + CGC — 31.53 (2019-10-12) NASNet-A — 35.77 (2019-12-01) PNASNet-5L — 35.63 (2019-12-01) Inception-v4 — 32.24 (2019-12-01) ResNet-152 — 29.59 (2019-12-01) VGG-14 — 19.13 (2019-12-01) AlexNet — 6.78 (2019-12-01) BiT-L (ResNet-152x4) — 58.7 (2019-12-24) BiT-M (ResNet-152x4) — 47.0 (2019-12-24) BiT-S (ResNet-152x4) — 36.0 (2019-12-24) ResNet-50 + MixUp (rescaled) — 28.37 (2020-06-10) ResNet-50 + GroupNorm — 29.2 (2020-06-30) ResNet-50 + RoHL — 29.2 (2020-06-30) ResNet-50 + FixUp — 28.5 (2020-06-30) BigBiGAN (RevNet-50 4×) — 4.92 (2020-08-24) ResNet-152 (FRCNN-ag-ad, VOC) — 13.2 (2020-11-28) ResNet-152 + GenInt with Transfer — 39.38 (2020-12-22) ResNet-18 + GenInt with Transfer — 27.03 (2020-12-22) CLIP — 72.3 (2021-02-26) BYOL (BG_RM) — 23.9 (2021-03-23) SwAV (BG_RM) — 21.9 (2021-03-23) MoCo-v2 (BG_Swaps) — 20.8 (2021-03-23) ViT-G/14 — 70.53 (2021-06-08) NS (Eff.-L2) — 68.5 (2021-06-08) ResNet34-RPG — 16.5 (2021-07-15) ViT-B/16 (ANN-1.3B) — 50.7 (2021-08-12) ResNet-101 (JFT-300M) — 49.1 (2021-08-12) ViT-B/32 — 48.4 (2021-08-12) ResNet-50 (JFT-300M) — 42.5 (2021-08-12) WiSE-FT — 72.1 (2021-09-04) C-BYOL — 25.5 (2021-09-27) C-SimCLR — 20.8 (2021-09-27) SeLa(v2) (reverse linear probing) — 20.61 (2021-09-29) DeepCluster(v2) (reverse linear probing) — 19.73 (2021-09-29) SwAV (reverse linear probing) — 17.71 (2021-09-29) MoCo(v2) (reverse linear probing) — 12.67 (2021-09-29) MoCHi (reverse linear probing) — 12.64 (2021-09-29) OBoW (reverse linear probing) — 12.23 (2021-09-29) LiT — 82.5 (2021-11-15) BASIC — 82.3 (2021-11-19) ALIGN — 72.2 (2021-11-19) ViT-B (Discrete 512x512) — 46.62 (2021-11-20) ViT-B/16 (512x512) + Pyramid — 49.39 (2021-11-30) ViT-B/16 (512x512) + Pixel — 47.53 (2021-11-30) ViT-B/16 (512x512) — 46.68 (2021-11-30) RegViT on 384x384 + Adv Pyramid — 39.79 (2021-11-30) RegViT on 384x384 + Adv Pixel — 37.41 (2021-11-30) RegViT on 384x384 — 35.59 (2021-11-30) RegViT on 384x384 + Random Pyramid — 34.83 (2021-11-30) RegViT on 384x384 + Random Pixel — 34.12 (2021-11-30) RegViT (RandAug) + Adv Pyramid — 32.92 (2021-11-30) Discrete ViT + Pixel — 30.98 (2021-11-30) Discrete ViT + Pyramid — 30.28 (2021-11-30) RegViT (RandAug) + Adv Pixel — 30.11 (2021-11-30) Discrete ViT — 29.95 (2021-11-30) RegViT (RandAug) + Random Pyramid — 29.41 (2021-11-30) RegViT (RandAug) — 29.3 (2021-11-30) RegViT (RandAug) + Random Pixel — 28.72 (2021-11-30) MLP-Mixer + Pyramid — 28.6 (2021-11-30) MLP-Mixer — 25.9 (2021-11-30) ViT + MixUp — 25.65 (2021-11-30) MLP-Mixer + Pixel — 24.75 (2021-11-30) ViT + CutMix — 21.61 (2021-11-30) ViT — 17.36 (2021-11-30) CLIP L — 42.8 (2021-12-31) CLIP L (LAION) — 42.1 (2021-12-31) RELICv2 — 25.9 (2022-01-13) RELIC — 23.8 (2022-01-13) BYOL — 23.0 (2022-01-13) SimCLR — 14.6 (2022-01-13) SWAG (ViT H/14) — 69.5 (2022-01-20) RegNetY 128GF (Platt) — 64.3 (2022-01-20) ViT H/14 (Platt) — 60.0 (2022-01-20) ViT L/16 (Platt) — 57.3 (2022-01-20) ViT B/16 — 48.9 (2022-01-20) SEER (RegNet10B) — 60.2 (2022-02-16) Baseline (ViT-G/14) — 79.03 (2022-03-10) Model soups (ViT-G/14) — 78.52 (2022-03-10) Vit B/16 (Bamboo) — 53.9 (2022-03-15) ResNet-50 (Bamboo) — 38.8 (2022-03-15) DILEMMA — 20.51 (2022-04-10) CLIP (CC12M pretrain) — 15.24 (2022-04-10) ResNet-50 (ImageNet-Captions) — 18.7 (2022-05-03) CoCa — 82.7 (2022-05-04) ALIGN-MRL — 51.6 (2022-05-26) AR-L (Opt Relevance) — 52.0 (2022-06-02) AR-B (Opt Relevance) — 47.1 (2022-06-02) AR-L — 46.5 (2022-06-02) ViT-L (Opt Relevance) — 43.2 (2022-06-02) ViT-B (Opt Relevance) — 42.2 (2022-06-02) AR-B — 41.4 (2022-06-02) AR-S (Opt Relevance) — 39.3 (2022-06-02) ViT-L — 37.4 (2022-06-02) DeiT-L (Opt Relevance) — 36.3 (2022-06-02) ViT-B — 35.1 (2022-06-02) AR-S — 34.3 (2022-06-02) DeiT-S (Opt Relevance) — 31.6 (2022-06-02) DeiT-L — 31.4 (2022-06-02) DeiT-S — 28.3 (2022-06-02) ViT-e — 72.0 (2022-09-14) LLE (ViT-H/14, MAE, Edge Aug) — 60.78 (2022-12-09) MAWS (ViT-6.5B) — 77.9 (2023-03-23) MAWS (ViT-2B) — 75.8 (2023-03-23) MAWS (ViT-H) — 72.6 (2023-03-23) EVA-02-CLIP-E/14+ — 79.6 (2023-03-27) ResNet-50 + CGC — 31.53 (2019-10-12) NASNet-A — 35.77 (2019-12-01) BiT-L (ResNet-152x4) — 58.7 (2019-12-24) CLIP — 72.3 (2021-02-26) LiT — 82.5 (2021-11-15) CoCa — 82.7 (2022-05-04)
RankModel Top-1 AccuracyTop-5 Accuracy Extra Training Data PaperCodeYear
1 CoCa 82.7 CoCa: Contrastive Captioners are Image-Text Foundation Models mlfoundations/open_clip · facebookresearch/multimodal · lucidrains/CoCa-pytorch · +3 2022
2 LiT 82.5 LiT: Zero-Shot Transfer with Locked-image text Tuning mlfoundations/open_clip · google-research/vision_transformer · google-research/big_vision · +2 2021
3 BASIC 82.3 Combined Scaling for Zero-shot Transfer Learning 2021
4 EVA-02-CLIP-E/14+ 79.6 EVA-CLIP: Improved Training Techniques for CLIP at Scale baaivision/eva · PaddlePaddle/PaddleMIX · Yui010206/CREMA · +1 2023
5 Baseline (ViT-G/14) 79.03 Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time mlfoundations/model-soups · Burf/ModelSoups · facebookresearch/ModelRatatouille · +3 2022
6 Model soups (ViT-G/14) 78.52 Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time mlfoundations/model-soups · Burf/ModelSoups · facebookresearch/ModelRatatouille · +3 2022
7 MAWS (ViT-6.5B) 77.9 The effectiveness of MAE pre-pretraining for billion-scale pretraining facebookresearch/maws 2023
8 MAWS (ViT-2B) 75.8 The effectiveness of MAE pre-pretraining for billion-scale pretraining facebookresearch/maws 2023
9 MAWS (ViT-H) 72.6 The effectiveness of MAE pre-pretraining for billion-scale pretraining facebookresearch/maws 2023
10 CLIP 72.3 Learning Transferable Visual Models From Natural Language Supervision openai/CLIP · mlfoundations/open_clip · towhee-io/towhee · +79 2021
11 ALIGN 72.2 Combined Scaling for Zero-shot Transfer Learning 2021
12 WiSE-FT 72.1 Robust fine-tuning of zero-shot models mlfoundations/wise-ft · mlfoundations/model-soups · ivanaer/g-universal-clip 2021
13 ViT-e 72.0 PaLI: A Jointly-Scaled Multilingual Language-Image Model google-research/big_vision 2022
14 ViT-G/14 70.53 Scaling Vision Transformers google-research/big_vision 2021
15 SWAG (ViT H/14) 69.5 Revisiting Weakly Supervised Pre-Training of Visual Perception Models facebookresearch/SWAG · Expedit-LargeScale-Vision-Transformer/Expedit-SWAG 2022
16 NS (Eff.-L2) 68.5 Scaling Vision Transformers google-research/big_vision 2021
17 RegNetY 128GF (Platt) 64.3 Revisiting Weakly Supervised Pre-Training of Visual Perception Models facebookresearch/SWAG · Expedit-LargeScale-Vision-Transformer/Expedit-SWAG 2022
18 LLE (ViT-H/14, MAE, Edge Aug) 60.78 A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others facebookresearch/Whac-A-Mole 2022
19 SEER (RegNet10B) 60.2 Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision facebookresearch/vissl 2022
20 ViT H/14 (Platt) 60 Revisiting Weakly Supervised Pre-Training of Visual Perception Models facebookresearch/SWAG · Expedit-LargeScale-Vision-Transformer/Expedit-SWAG 2022
1–20 / 106 다음 → 페이지당 10 20 50 100