| Rank | Model |
box AP | AP50 | AP75 | APS | APM | APL | Params (M) |
Extra Training Data |
Paper | Code | Year |
| 1 |
PE_spatial (DETA) |
66.0 | – | – | – | – | – | 1900 |
✓ |
Perception Encoder: The best visual embeddings are not at the output of the network
|
facebookresearch/perception_models |
2025 |
| 2 |
Co-DETR |
65.9 | – | – | – | – | – | 314 |
✓ |
DETRs with Collaborative Hybrid Assignments Training
|
open-mmlab/mmdetection · siyuanliii/masa · sense-x/co-detr
· +3 |
2022 |
| 3 |
M3I Pre-training (InternImage-H) |
65.0 | – | – | – | – | – | – |
✓ |
Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information
|
OpenGVLab/M3I-Pretraining |
2022 |
| 3 |
InternImage-H |
65.0 | – | – | – | – | – | – |
✓ |
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
|
opengvlab/internimage · OpenGVLab/M3I-Pretraining · chenller/mmseg-extension |
2022 |
| 5 |
Co-DETR (Swin-L) |
64.7 | – | – | – | – | – | 218 |
✓ |
DETRs with Collaborative Hybrid Assignments Training
|
open-mmlab/mmdetection · siyuanliii/masa · sense-x/co-detr
· +3 |
2022 |
| 6 |
Focal-Stable-DINO (Focal-Huge, no TTA) |
64.6 | 81.5 | 71.4 | 50.4 | 68.5 | 78.5 | 689 |
✓ |
A Strong and Reproducible Object Detector with Only Public Datasets
|
microsoft/FocalNet · idea-research/stable-dino · idea-research/stabledino |
2023 |
| 7 |
EVA |
64.5 | 82.1 | 70.8 | 49.4 | 68.4 | 78.5 | – |
✓ |
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
|
rwightman/pytorch-image-models · open-mmlab/mmselfsup · baaivision/eva
· +3 |
2022 |
| 8 |
ViT-CoMer |
64.3 | – | – | – | – | – | 363 |
|
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
|
Traffic-X/ViT-CoMer · chenller/mmseg-extension |
2024 |
| 9 |
FocalNet-H (DINO) |
64.2 | – | – | – | – | – | – |
✓ |
Focal Modulation Networks
|
PaddlePaddle/PaddleDetection · keras-team/keras-io · microsoft/FocalNet
· +6 |
2022 |
| 9 |
InternImage-XL |
64.2 | – | – | – | – | – | – |
✓ |
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
|
opengvlab/internimage · OpenGVLab/M3I-Pretraining · chenller/mmseg-extension |
2022 |
| 11 |
CP-DETR-L Swin-L(Fine tuning separately in COCO) |
64.1 | – | – | – | – | – | – |
✓ |
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection
|
|
2024 |
| 12 |
RevCol-H(DINO) |
63.8 | – | – | – | – | – | – |
✓ |
Reversible Column Networks
|
megvii-research/revcol |
2022 |
| 13 |
DINO (Swin-L) |
63.2 | – | – | – | – | – | – |
|
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
|
IDEA-Research/Grounded-Segment-Anything · PaddlePaddle/PaddleDetection · lucasjinreal/yolov7_d2
· +13 |
2022 |
| 14 |
Grounding DINO |
63.0 | – | – | – | – | – | – |
✓ |
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
|
huggingface/transformers · IDEA-Research/Grounded-Segment-Anything · idea-research/groundingdino
· +7 |
2023 |
| 15 |
SwinV2-G (HTC++) |
62.5 | – | – | – | – | – | – |
✓ |
Swin Transformer V2: Scaling Up Capacity and Resolution
|
rwightman/pytorch-image-models · microsoft/Swin-Transformer · PaddlePaddle/PaddleDetection
· +20 |
2021 |
| 16 |
Florence-CoSwin-H |
62 | – | – | – | – | – | – |
✓ |
Florence: A New Foundation Model for Computer Vision
|
microsoft/unicl · MindCode-4/code-3 |
2021 |
| 16 |
GLEE-Pro |
62.0 | – | – | – | – | – | – |
✓ |
General Object Foundation Model for Images and Videos at Scale
|
FoundationVision/GLEE |
2023 |
| 18 |
ViTDet, ViT-H Cascade (multiscale) |
61.3 | – | – | – | – | – | – |
|
Exploring Plain Vision Transformer Backbones for Object Detection
|
facebookresearch/detectron2 · PaddlePaddle/PaddleDetection · alibaba/EasyCV
· +8 |
2022 |
| 19 |
GLIP (Swin-L, multi-scale) |
60.8 | – | – | – | – | – | – |
✓ |
Grounded Language-Image Pre-training
|
microsoft/GLIP · brown-palm/ObjectPrompt · rsCPSyEu/ovd_cod |
2021 |
| 20 |
Soft Teacher + Swin-L (HTC++, multi-scale) |
60.7 | – | – | – | – | – | – |
✓ |
End-to-End Semi-Supervised Object Detection with Soft Teacher
|
microsoft/SoftTeacher · amazon-science/bigdetection · amazon-research/bigdetection
· +5 |
2021 |
| 21 |
UNINEXT-H |
60.6 | 77.5 | 66.7 | 45.1 | 64.8 | 75.3 | – |
✓ |
Universal Instance Perception as Object Discovery and Retrieval
|
MasterBin-IIAU/UNINEXT |
2023 |
| 22 |
ViT-Adapter-L (HTC++, BEiTv2 pretrain, multi-scale) |
60.5 | – | – | – | – | – | – |
|
Vision Transformer Adapter for Dense Predictions
|
czczup/vit-adapter · chenller/mmseg-extension |
2022 |
| 23 |
ViTDet, ViT-H Cascade |
60.4 | – | – | – | – | – | – |
|
Exploring Plain Vision Transformer Backbones for Object Detection
|
facebookresearch/detectron2 · PaddlePaddle/PaddleDetection · alibaba/EasyCV
· +8 |
2022 |
| 23 |
GLEE-Plus |
60.4 | – | – | – | – | – | – |
✓ |
General Object Foundation Model for Images and Videos at Scale
|
FoundationVision/GLEE |
2023 |
| 25 |
DyHead (Swin-L, multi scale, self-training) |
60.3 | 78.2 | – | – | – | 74.2 | – |
✓ |
Dynamic Head: Unifying Object Detection Heads with Attentions
|
open-mmlab/mmdetection · microsoft/DynamicHead · Coldestadam/DynamicHead |
2021 |
| 26 |
ViT-Adapter-L (HTC++, BEiT pretrain, multi-scale) |
60.2 | – | – | – | – | – | – |
|
Vision Transformer Adapter for Dense Predictions
|
czczup/vit-adapter · chenller/mmseg-extension |
2022 |
| 27 |
Soft Teacher+Swin-L(HTC++, single scale) |
60.1 | – | – | – | – | – | – |
✓ |
End-to-End Semi-Supervised Object Detection with Soft Teacher
|
microsoft/SoftTeacher · amazon-science/bigdetection · amazon-research/bigdetection
· +5 |
2021 |
| 28 |
CBNetV2 (Dual-Swin-L HTC, multi-scale) |
59.6 | – | – | – | – | – | – |
|
CBNet: A Composite Backbone Network Architecture for Object Detection
|
PaddlePaddle/PaddleDetection · shinya7y/UniverseNet · VDIGPKU/CBNetV2
· +1 |
2021 |
| 29 |
Frozen Backbone, SwinV2-G-ext22K (HTC) |
59.3 | – | – | – | – | – | – |
|
Could Giant Pretrained Image Models Extract Universal Representations?
|
|
2022 |
| 30 |
HorNet-L |
59.2 | – | – | – | – | – | – |
|
HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions
|
open-mmlab/mmclassification · towhee-io/towhee · chengtan9907/OpenSTL
· +5 |
2022 |
| 30 |
MOAT-3 (IN-22K pretraining, single-scale) |
59.2 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 32 |
CBNetV2 (Dual-Swin-L HTC, multi-scale) |
59.1 | – | – | – | – | – | – |
|
CBNet: A Composite Backbone Network Architecture for Object Detection
|
PaddlePaddle/PaddleDetection · shinya7y/UniverseNet · VDIGPKU/CBNetV2
· +1 |
2021 |
| 33 |
Focal-L (DyHead, multi-scale) |
58.7 | 77.2 | – | – | – | 73.4 | – |
|
Focal Self-attention for Local-Global Interactions in Vision Transformers
|
BR-IDL/PaddleViT · microsoft/Focal-Transformer · microsoft/esvit |
2021 |
| 33 |
MViTv2-L (Cascade Mask R-CNN, multi-scale, IN21k pre-train) |
58.7 | – | – | – | – | – | – |
|
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
|
rwightman/pytorch-image-models · facebookresearch/detectron2 · facebookresearch/SlowFast
· +6 |
2021 |
| 35 |
MOAT-2 (IN-22K pretraining, single-scale) |
58.5 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 36 |
DyHead (Swin-L, multi scale) |
58.4 | 76.8 | – | 44.5 | 62.2 | 73.2 | – |
|
Dynamic Head: Unifying Object Detection Heads with Attentions
|
open-mmlab/mmdetection · microsoft/DynamicHead · Coldestadam/DynamicHead |
2021 |
| 37 |
Swin-L (HTC++, multi scale) |
58 | – | – | – | – | – | – |
|
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
|
huggingface/transformers · rwightman/pytorch-image-models · open-mmlab/mmdetection
· +77 |
2021 |
| 38 |
MOAT-1 (IN-1K pretraining, single-scale) |
57.7 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 39 |
UM-MAE(HTC++, Swin-L, IN1K) |
57.4 | – | – | – | – | – | – |
|
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality
|
implus/um-mae |
2022 |
| 40 |
YOLOv6-L6(46 fps, 1280, V100) |
57.2 | 74.5 | – | – | – | – | – |
|
YOLOv6 v3.0: A Full-Scale Reloading
|
PaddlePaddle/PaddleDetection · meituan/yolov6 · PaddlePaddle/PaddleYOLO
· +2 |
2023 |
| 41 |
Swin-L (HTC++, single scale) |
57.1 | – | – | – | – | – | – |
|
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
|
huggingface/transformers · rwightman/pytorch-image-models · open-mmlab/mmdetection
· +77 |
2021 |
| 41 |
TransNeXt-Base (IN-1K pretrain, DINO 1x) |
57.1 | – | – | – | – | – | – |
|
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
|
Westlake-AI/openmixup · daishiresearch/transnext · chenller/mmseg-extension
· +1 |
2023 |
| 43 |
Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale) |
57.0 | – | – | – | – | – | – |
✓ |
Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation
|
PaddlePaddle/PaddleOCR · open-mmlab/mmdetection · tensorflow/tpu
· +2 |
2020 |
| 44 |
TransNeXt-Small (IN-1K pretrain, DINO 1x) |
56.6 | – | – | – | – | – | – |
|
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
|
Westlake-AI/openmixup · daishiresearch/transnext · chenller/mmseg-extension
· +1 |
2023 |
| 45 |
QueryInst (single scale) |
56.1 | 75.8 | 61.7 | 40.2 | 59.8 | 71.5 | – |
|
Instances as Queries
|
open-mmlab/mmdetection · hustvl/QueryInst · Bo396543018/picodet_repro
· +2 |
2021 |
| 45 |
MViTv2-H (Cascade Mask R-CNN, single-scale, IN21k pre-train) |
56.1 | – | – | – | – | – | – |
|
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
|
rwightman/pytorch-image-models · facebookresearch/detectron2 · facebookresearch/SlowFast
· +6 |
2021 |
| 47 |
MOAT-0 (IN-1K pretraining, single-scale) |
55.9 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 48 |
TransNeXt-Tiny (IN-1K pretrain, DINO 1x) |
55.7 | – | – | – | – | – | – |
|
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
|
Westlake-AI/openmixup · daishiresearch/transnext · chenller/mmseg-extension
· +1 |
2023 |
| 49 |
YOLOv4-P7 CSP-P7 (single-scale, 16 fps) |
55.4 | 73.3 | 60.7 | 38.1 | 59.5 | 67.4 | – |
|
Scaled-YOLOv4: Scaling Cross Stage Partial Network
|
AlexeyAB/darknet · RangiLyu/nanodet · WongKinYiu/ScaledYOLOv4
· +38 |
2020 |
| 50 |
tiny-MOAT-3 (IN-1K pretraining, single-scale) |
55.2 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |