| Rank | Model |
box AP | AP50 | AP75 | APS | APM | APL | Params (M) |
Extra Training Data |
Paper | Code | Year |
| 21 |
UNINEXT-H |
60.6 | 77.5 | 66.7 | 45.1 | 64.8 | 75.3 | – |
✓ |
Universal Instance Perception as Object Discovery and Retrieval
|
MasterBin-IIAU/UNINEXT |
2023 |
| 22 |
ViT-Adapter-L (HTC++, BEiTv2 pretrain, multi-scale) |
60.5 | – | – | – | – | – | – |
|
Vision Transformer Adapter for Dense Predictions
|
czczup/vit-adapter · chenller/mmseg-extension |
2022 |
| 23 |
ViTDet, ViT-H Cascade |
60.4 | – | – | – | – | – | – |
|
Exploring Plain Vision Transformer Backbones for Object Detection
|
facebookresearch/detectron2 · PaddlePaddle/PaddleDetection · alibaba/EasyCV
· +8 |
2022 |
| 23 |
GLEE-Plus |
60.4 | – | – | – | – | – | – |
✓ |
General Object Foundation Model for Images and Videos at Scale
|
FoundationVision/GLEE |
2023 |
| 25 |
DyHead (Swin-L, multi scale, self-training) |
60.3 | 78.2 | – | – | – | 74.2 | – |
✓ |
Dynamic Head: Unifying Object Detection Heads with Attentions
|
open-mmlab/mmdetection · microsoft/DynamicHead · Coldestadam/DynamicHead |
2021 |
| 26 |
ViT-Adapter-L (HTC++, BEiT pretrain, multi-scale) |
60.2 | – | – | – | – | – | – |
|
Vision Transformer Adapter for Dense Predictions
|
czczup/vit-adapter · chenller/mmseg-extension |
2022 |
| 27 |
Soft Teacher+Swin-L(HTC++, single scale) |
60.1 | – | – | – | – | – | – |
✓ |
End-to-End Semi-Supervised Object Detection with Soft Teacher
|
microsoft/SoftTeacher · amazon-science/bigdetection · amazon-research/bigdetection
· +5 |
2021 |
| 28 |
CBNetV2 (Dual-Swin-L HTC, multi-scale) |
59.6 | – | – | – | – | – | – |
|
CBNet: A Composite Backbone Network Architecture for Object Detection
|
PaddlePaddle/PaddleDetection · shinya7y/UniverseNet · VDIGPKU/CBNetV2
· +1 |
2021 |
| 29 |
Frozen Backbone, SwinV2-G-ext22K (HTC) |
59.3 | – | – | – | – | – | – |
|
Could Giant Pretrained Image Models Extract Universal Representations?
|
|
2022 |
| 30 |
HorNet-L |
59.2 | – | – | – | – | – | – |
|
HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions
|
open-mmlab/mmclassification · towhee-io/towhee · chengtan9907/OpenSTL
· +5 |
2022 |
| 30 |
MOAT-3 (IN-22K pretraining, single-scale) |
59.2 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 32 |
CBNetV2 (Dual-Swin-L HTC, multi-scale) |
59.1 | – | – | – | – | – | – |
|
CBNet: A Composite Backbone Network Architecture for Object Detection
|
PaddlePaddle/PaddleDetection · shinya7y/UniverseNet · VDIGPKU/CBNetV2
· +1 |
2021 |
| 33 |
Focal-L (DyHead, multi-scale) |
58.7 | 77.2 | – | – | – | 73.4 | – |
|
Focal Self-attention for Local-Global Interactions in Vision Transformers
|
BR-IDL/PaddleViT · microsoft/Focal-Transformer · microsoft/esvit |
2021 |
| 33 |
MViTv2-L (Cascade Mask R-CNN, multi-scale, IN21k pre-train) |
58.7 | – | – | – | – | – | – |
|
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
|
rwightman/pytorch-image-models · facebookresearch/detectron2 · facebookresearch/SlowFast
· +6 |
2021 |
| 35 |
MOAT-2 (IN-22K pretraining, single-scale) |
58.5 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 36 |
DyHead (Swin-L, multi scale) |
58.4 | 76.8 | – | 44.5 | 62.2 | 73.2 | – |
|
Dynamic Head: Unifying Object Detection Heads with Attentions
|
open-mmlab/mmdetection · microsoft/DynamicHead · Coldestadam/DynamicHead |
2021 |
| 37 |
Swin-L (HTC++, multi scale) |
58 | – | – | – | – | – | – |
|
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
|
huggingface/transformers · rwightman/pytorch-image-models · open-mmlab/mmdetection
· +77 |
2021 |
| 38 |
MOAT-1 (IN-1K pretraining, single-scale) |
57.7 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 39 |
UM-MAE(HTC++, Swin-L, IN1K) |
57.4 | – | – | – | – | – | – |
|
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality
|
implus/um-mae |
2022 |
| 40 |
YOLOv6-L6(46 fps, 1280, V100) |
57.2 | 74.5 | – | – | – | – | – |
|
YOLOv6 v3.0: A Full-Scale Reloading
|
PaddlePaddle/PaddleDetection · meituan/yolov6 · PaddlePaddle/PaddleYOLO
· +2 |
2023 |