| Rank | Model |
MAP | AP50 | mAP@50 | mAP@50-95 |
Extra Training Data |
Paper | Code | Year |
| 1 |
Cascade Eff-B7 NAS-FPN (Copy Paste pre-training, single-scale) |
89.3% | – | – | – |
|
Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation
|
PaddlePaddle/PaddleOCR · open-mmlab/mmdetection · tensorflow/tpu
· +2 |
2020 |
| 2 |
YOLO-Former |
86.01% | – | – | – |
✓ |
YOLO-Former: YOLO Shakes Hand With ViT
|
|
2024 |
| 3 |
DETReg (MDef-DETR) |
84.16% | 84.16 | – | – |
✓ |
Class-agnostic Object Detection with Multi-modal Transformer
|
mmaaz60/mvits_for_class_agnostic_od |
2021 |
| 4 |
HSD (VGG16, 512x512, single-scale test) |
83.0% | – | – | – |
|
Hierarchical Shot Detector
|
JialeCao001/HSD |
2019 |
| 5 |
CoupleNet |
82.7% | – | – | – |
|
CoupleNet: Coupling Global Structure with Local Parts for Object Detection
|
princewang1994/R-FCN.pytorch · tshizys/CoupleNet · princewang1994/RFCN_CoupleNet.pytorch |
2017 |
| 6 |
EEEA-Net-C2 (YOLOv4) |
81.8% | – | – | – |
✓ |
EEEA-Net: An Early Exit Evolutionary Neural Architecture Search
|
chakkritte/eeea-net |
2021 |
| 7 |
HSD (VGG16, 320x320, single-scale test) |
81.7% | – | – | – |
|
Hierarchical Shot Detector
|
JialeCao001/HSD |
2019 |
| 8 |
SSD512 (07+12+COCO) |
81.6% | – | – | – |
✓ |
SSD: Single Shot MultiBox Detector
|
open-mmlab/mmdetection · serengil/deepface · pytorch/vision
· +218 |
2015 |
| 9 |
BlitzNet512 + seg (s8) |
81.5% | – | – | – |
|
BlitzNet: A Real-Time Deep Network for Scene Understanding
|
dvornikita/blitznet · ShunyuYao/blitznet_instance_segment |
2017 |
| 9 |
Localize |
81.5% | – | – | – |
|
Localize to Classify and Classify to Localize: Mutual Guidance in Object Detection
|
ZHANGHeng19931123/MutualGuide |
2020 |
| 11 |
CenterNet(DLA34, Flip, 512x512) |
80.7% | – | – | – |
|
Objects as Points
|
tensorflow/models · open-mmlab/mmdetection · PaddlePaddle/PaddleDetection
· +73 |
2019 |
| 12 |
PS-KD (ResNet-152, CutMix) |
79.7% | – | – | – |
|
Self-Knowledge Distillation with Progressive Refinement of Targets
|
lgcnsai/ps-kd-pytorch |
2020 |
| 13 |
DPNet |
79.2% | – | – | – |
|
DPNet: Dual-Path Network for Real-time Object Detection with Lightweight Attention
|
huiminshii/dpnet · MS-Mind/MS-Code-02 |
2022 |
| 14 |
OHEM |
78.9% | – | – | – |
|
Training Region-based Object Detectors with Online Hard Example Mining
|
abhi2610/ohem · tkuanlun350/Kaggle_Ship_Detection_2018 · Bennie-Han/Image-augementation-pytorch
· +2 |
2016 |
| 15 |
YOLO v2 |
78.6% | – | – | – |
✓ |
YOLO9000: Better, Faster, Stronger
|
AlexeyAB/darknet · PaddlePaddle/PaddleDetection · thtrieu/darkflow
· +228 |
2016 |
| 15 |
ThunderNet SNet535 Backbone |
78.6% | – | – | – |
|
ThunderNet: Towards Real-time Generic Object Detection
|
ouyanghuiyu/Thundernet_Pytorch · qinzheng93/thundernet · saswat0/Thundernet-Object-Detection |
2019 |
| 17 |
DeNet-101 (skip) |
77.1% | – | – | – |
|
DeNet: Scalable Real-time Object Detection with Directed Sparse Sampling
|
lachlants/denet |
2017 |
| 18 |
I+ORE |
76.2% | – | – | – |
|
Random Erasing Data Augmentation
|
rwightman/pytorch-image-models · pytorch/vision · albumentations-team/albumentations
· +15 |
2017 |
| 19 |
Perona Malik (Perona and Malik, 1990) |
74.37% | – | – | – |
|
Learning Visual Representations for Transfer Learning by Suppressing Texture
|
HaohanWang/ImageNet-Sketch |
2020 |
| 20 |
FRCN |
74.2% | – | – | – |
|
A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
|
xiaolonw/adversarial-frcnn · HusterRC/adversarial-frcnn-master · busyboxs/Some-resources-useful-for-me
· +1 |
2017 |