paper-with-me

Object Detection 벤치마크

Object Detection on COCO 2017

120개 결과 · ⬇ CSV · JSON

AP

40.7 43.88 47.05 50.23 53.4 2022-04 2026-09 MaxViT-B — 53.4 (2022-04-04) MaxViT-S — 53.1 (2022-04-04) MaxViT-T — 52.1 (2022-04-04) MaxViT-B — 53.4 (2022-04-04) MaxViT-S — 53.1 (2022-04-04) MaxViT-T — 52.1 (2022-04-04) MaxViT-B — 53.4 (2022-04-04) MaxViT-S — 53.1 (2022-04-04) MaxViT-T — 52.1 (2022-04-04) MaxViT-B — 53.4 (2022-04-04) MaxViT-S — 53.1 (2022-04-04) MaxViT-T — 52.1 (2022-04-04) MaxViT-B — 53.4 (2022-04-04) MaxViT-S — 53.1 (2022-04-04) MaxViT-T — 52.1 (2022-04-04) Faster R-CNN (ideal number of groups) — 40.7 (2023-02-07) Faster R-CNN (ideal number of groups) — 40.7 (2023-02-07) Faster R-CNN (ideal number of groups) — 40.7 (2023-02-07) Faster R-CNN (ideal number of groups) — 40.7 (2023-02-07) Faster R-CNN (ideal number of groups) — 40.7 (2023-02-07) DAT-S++ — 50.2 (2023-09-04) DAT-T++ — 49.2 (2023-09-04) DAT-S++ — 50.2 (2023-09-04) DAT-T++ — 49.2 (2023-09-04) DAT-S++ — 50.2 (2023-09-04) DAT-T++ — 49.2 (2023-09-04) DAT-S++ — 50.2 (2023-09-04) DAT-T++ — 49.2 (2023-09-04) DAT-S++ — 50.2 (2023-09-04) DAT-T++ — 49.2 (2023-09-04) DyHead (SAP) — 42.1 (2024-09-25) DyHead (SAP) — 42.1 (2024-09-25) DyHead (SAP) — 42.1 (2024-09-25) DyHead (SAP) — 42.1 (2024-09-25) DyHead (SAP) — 42.1 (2024-09-25) MaxViT-B — 53.4 (2022-04-04)
RankModel APmAPMean mAPAP50AP75APMAPM50APM75 PaperCodeYear
1 MaxViT-B 53.472.958.145.770.350 MaxViT: Multi-Axis Vision Transformer huggingface/pytorch-image-models · lucidrains/vit-pytorch · lucidrains/imagen-pytorch · +12 2022
2 MaxViT-S 53.172.558.145.469.849.5 MaxViT: Multi-Axis Vision Transformer huggingface/pytorch-image-models · lucidrains/vit-pytorch · lucidrains/imagen-pytorch · +12 2022
3 MaxViT-T 52.171.956.844.669.148.4 MaxViT: Multi-Axis Vision Transformer huggingface/pytorch-image-models · lucidrains/vit-pytorch · lucidrains/imagen-pytorch · +12 2022
4 DAT-S++ 50.2 DAT++: Spatially Dynamic Vision Transformer with Deformable Attention leaplabthu/dat 2023
5 DAT-T++ 49.2 DAT++: Spatially Dynamic Vision Transformer with Deformable Attention leaplabthu/dat 2023
6 DyHead (SAP) 42.159.445.9 Stochastic Subsampling With Average Pooling 2024
7 Faster R-CNN (ideal number of groups) 40.761.244.6 On the Ideal Number of Groups for Isometric Gradient Propagation 2023
8 UniRepLKNet-XL++ 56.4 UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension 2023
9 UniRepLKNet-L++ 55.8 UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension 2023
10 UniRepLKNet-B++ 54.8 UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension 2023
11 UniRepLKNet-S++ 54.3 UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension 2023
12 MixMIM-L 54.1 MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers sense-x/mixmim 2022
13 UniRepLKNet-S 53 UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension 2023
14 MixMIM-B 52.2 MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers sense-x/mixmim 2022
15 UniRepLKNet-T 51.7 UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension 2023
16 BiFormer-B (IN1k pretrain, MaskRCNN 12ep) 48.6 BiFormer: Vision Transformer with Bi-Level Routing Attention rayleizhu/biformer · chenller/mmseg-extension · birder/birder 2023
17 DeBiFormer-B (IN1k pretrain, MaskRCNN 12ep) 48.5 DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention maclong01/DeBiFormer 2024
18 BiFormer-S (IN1k pretrain, MaskRCNN 12ep) 47.8 BiFormer: Vision Transformer with Bi-Level Routing Attention rayleizhu/biformer · chenller/mmseg-extension · birder/birder 2023
19 DeBiFormer-S (IN1k pretrain, MaskRCNN 12ep) 47.5 DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention maclong01/DeBiFormer 2024
20 DeBiFormer-B (IN1k pretrain, Retina) 47.1 DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention maclong01/DeBiFormer 2024
1–20 / 120 다음 → 페이지당 10 20 50 100