| Rank | Model |
box AP | AP50 | AP75 | APS | APM | APL | Params (M) |
Extra Training Data |
Paper | Code | Year |
| 1 |
PE_spatial (DETA) |
66.0 | – | – | – | – | – | 1900 |
✓ |
Perception Encoder: The best visual embeddings are not at the output of the network
|
facebookresearch/perception_models |
2025 |
| 2 |
Co-DETR |
65.9 | – | – | – | – | – | 314 |
✓ |
DETRs with Collaborative Hybrid Assignments Training
|
open-mmlab/mmdetection · siyuanliii/masa · sense-x/co-detr
· +3 |
2022 |
| 3 |
M3I Pre-training (InternImage-H) |
65.0 | – | – | – | – | – | – |
✓ |
Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information
|
OpenGVLab/M3I-Pretraining |
2022 |
| 3 |
InternImage-H |
65.0 | – | – | – | – | – | – |
✓ |
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
|
opengvlab/internimage · OpenGVLab/M3I-Pretraining · chenller/mmseg-extension |
2022 |
| 5 |
Co-DETR (Swin-L) |
64.7 | – | – | – | – | – | 218 |
✓ |
DETRs with Collaborative Hybrid Assignments Training
|
open-mmlab/mmdetection · siyuanliii/masa · sense-x/co-detr
· +3 |
2022 |
| 6 |
Focal-Stable-DINO (Focal-Huge, no TTA) |
64.6 | 81.5 | 71.4 | 50.4 | 68.5 | 78.5 | 689 |
✓ |
A Strong and Reproducible Object Detector with Only Public Datasets
|
microsoft/FocalNet · idea-research/stable-dino · idea-research/stabledino |
2023 |
| 7 |
EVA |
64.5 | 82.1 | 70.8 | 49.4 | 68.4 | 78.5 | – |
✓ |
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
|
rwightman/pytorch-image-models · open-mmlab/mmselfsup · baaivision/eva
· +3 |
2022 |
| 8 |
ViT-CoMer |
64.3 | – | – | – | – | – | 363 |
|
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
|
Traffic-X/ViT-CoMer · chenller/mmseg-extension |
2024 |
| 9 |
FocalNet-H (DINO) |
64.2 | – | – | – | – | – | – |
✓ |
Focal Modulation Networks
|
PaddlePaddle/PaddleDetection · keras-team/keras-io · microsoft/FocalNet
· +6 |
2022 |
| 9 |
InternImage-XL |
64.2 | – | – | – | – | – | – |
✓ |
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
|
opengvlab/internimage · OpenGVLab/M3I-Pretraining · chenller/mmseg-extension |
2022 |
| 11 |
CP-DETR-L Swin-L(Fine tuning separately in COCO) |
64.1 | – | – | – | – | – | – |
✓ |
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection
|
|
2024 |
| 12 |
RevCol-H(DINO) |
63.8 | – | – | – | – | – | – |
✓ |
Reversible Column Networks
|
megvii-research/revcol |
2022 |
| 13 |
DINO (Swin-L) |
63.2 | – | – | – | – | – | – |
|
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
|
IDEA-Research/Grounded-Segment-Anything · PaddlePaddle/PaddleDetection · lucasjinreal/yolov7_d2
· +13 |
2022 |
| 14 |
Grounding DINO |
63.0 | – | – | – | – | – | – |
✓ |
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
|
huggingface/transformers · IDEA-Research/Grounded-Segment-Anything · idea-research/groundingdino
· +7 |
2023 |
| 15 |
SwinV2-G (HTC++) |
62.5 | – | – | – | – | – | – |
✓ |
Swin Transformer V2: Scaling Up Capacity and Resolution
|
rwightman/pytorch-image-models · microsoft/Swin-Transformer · PaddlePaddle/PaddleDetection
· +20 |
2021 |
| 16 |
Florence-CoSwin-H |
62 | – | – | – | – | – | – |
✓ |
Florence: A New Foundation Model for Computer Vision
|
microsoft/unicl · MindCode-4/code-3 |
2021 |
| 16 |
GLEE-Pro |
62.0 | – | – | – | – | – | – |
✓ |
General Object Foundation Model for Images and Videos at Scale
|
FoundationVision/GLEE |
2023 |
| 18 |
ViTDet, ViT-H Cascade (multiscale) |
61.3 | – | – | – | – | – | – |
|
Exploring Plain Vision Transformer Backbones for Object Detection
|
facebookresearch/detectron2 · PaddlePaddle/PaddleDetection · alibaba/EasyCV
· +8 |
2022 |
| 19 |
GLIP (Swin-L, multi-scale) |
60.8 | – | – | – | – | – | – |
✓ |
Grounded Language-Image Pre-training
|
microsoft/GLIP · brown-palm/ObjectPrompt · rsCPSyEu/ovd_cod |
2021 |
| 20 |
Soft Teacher + Swin-L (HTC++, multi-scale) |
60.7 | – | – | – | – | – | – |
✓ |
End-to-End Semi-Supervised Object Detection with Soft Teacher
|
microsoft/SoftTeacher · amazon-science/bigdetection · amazon-research/bigdetection
· +5 |
2021 |
| 21 |
UNINEXT-H |
60.6 | 77.5 | 66.7 | 45.1 | 64.8 | 75.3 | – |
✓ |
Universal Instance Perception as Object Discovery and Retrieval
|
MasterBin-IIAU/UNINEXT |
2023 |
| 22 |
ViT-Adapter-L (HTC++, BEiTv2 pretrain, multi-scale) |
60.5 | – | – | – | – | – | – |
|
Vision Transformer Adapter for Dense Predictions
|
czczup/vit-adapter · chenller/mmseg-extension |
2022 |
| 23 |
ViTDet, ViT-H Cascade |
60.4 | – | – | – | – | – | – |
|
Exploring Plain Vision Transformer Backbones for Object Detection
|
facebookresearch/detectron2 · PaddlePaddle/PaddleDetection · alibaba/EasyCV
· +8 |
2022 |
| 23 |
GLEE-Plus |
60.4 | – | – | – | – | – | – |
✓ |
General Object Foundation Model for Images and Videos at Scale
|
FoundationVision/GLEE |
2023 |
| 25 |
DyHead (Swin-L, multi scale, self-training) |
60.3 | 78.2 | – | – | – | 74.2 | – |
✓ |
Dynamic Head: Unifying Object Detection Heads with Attentions
|
open-mmlab/mmdetection · microsoft/DynamicHead · Coldestadam/DynamicHead |
2021 |
| 26 |
ViT-Adapter-L (HTC++, BEiT pretrain, multi-scale) |
60.2 | – | – | – | – | – | – |
|
Vision Transformer Adapter for Dense Predictions
|
czczup/vit-adapter · chenller/mmseg-extension |
2022 |
| 27 |
Soft Teacher+Swin-L(HTC++, single scale) |
60.1 | – | – | – | – | – | – |
✓ |
End-to-End Semi-Supervised Object Detection with Soft Teacher
|
microsoft/SoftTeacher · amazon-science/bigdetection · amazon-research/bigdetection
· +5 |
2021 |
| 28 |
CBNetV2 (Dual-Swin-L HTC, multi-scale) |
59.6 | – | – | – | – | – | – |
|
CBNet: A Composite Backbone Network Architecture for Object Detection
|
PaddlePaddle/PaddleDetection · shinya7y/UniverseNet · VDIGPKU/CBNetV2
· +1 |
2021 |
| 29 |
Frozen Backbone, SwinV2-G-ext22K (HTC) |
59.3 | – | – | – | – | – | – |
|
Could Giant Pretrained Image Models Extract Universal Representations?
|
|
2022 |
| 30 |
HorNet-L |
59.2 | – | – | – | – | – | – |
|
HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions
|
open-mmlab/mmclassification · towhee-io/towhee · chengtan9907/OpenSTL
· +5 |
2022 |
| 30 |
MOAT-3 (IN-22K pretraining, single-scale) |
59.2 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 32 |
CBNetV2 (Dual-Swin-L HTC, multi-scale) |
59.1 | – | – | – | – | – | – |
|
CBNet: A Composite Backbone Network Architecture for Object Detection
|
PaddlePaddle/PaddleDetection · shinya7y/UniverseNet · VDIGPKU/CBNetV2
· +1 |
2021 |
| 33 |
Focal-L (DyHead, multi-scale) |
58.7 | 77.2 | – | – | – | 73.4 | – |
|
Focal Self-attention for Local-Global Interactions in Vision Transformers
|
BR-IDL/PaddleViT · microsoft/Focal-Transformer · microsoft/esvit |
2021 |
| 33 |
MViTv2-L (Cascade Mask R-CNN, multi-scale, IN21k pre-train) |
58.7 | – | – | – | – | – | – |
|
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
|
rwightman/pytorch-image-models · facebookresearch/detectron2 · facebookresearch/SlowFast
· +6 |
2021 |
| 35 |
MOAT-2 (IN-22K pretraining, single-scale) |
58.5 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 36 |
DyHead (Swin-L, multi scale) |
58.4 | 76.8 | – | 44.5 | 62.2 | 73.2 | – |
|
Dynamic Head: Unifying Object Detection Heads with Attentions
|
open-mmlab/mmdetection · microsoft/DynamicHead · Coldestadam/DynamicHead |
2021 |
| 37 |
Swin-L (HTC++, multi scale) |
58 | – | – | – | – | – | – |
|
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
|
huggingface/transformers · rwightman/pytorch-image-models · open-mmlab/mmdetection
· +77 |
2021 |
| 38 |
MOAT-1 (IN-1K pretraining, single-scale) |
57.7 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 39 |
UM-MAE(HTC++, Swin-L, IN1K) |
57.4 | – | – | – | – | – | – |
|
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality
|
implus/um-mae |
2022 |
| 40 |
YOLOv6-L6(46 fps, 1280, V100) |
57.2 | 74.5 | – | – | – | – | – |
|
YOLOv6 v3.0: A Full-Scale Reloading
|
PaddlePaddle/PaddleDetection · meituan/yolov6 · PaddlePaddle/PaddleYOLO
· +2 |
2023 |
| 41 |
Swin-L (HTC++, single scale) |
57.1 | – | – | – | – | – | – |
|
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
|
huggingface/transformers · rwightman/pytorch-image-models · open-mmlab/mmdetection
· +77 |
2021 |
| 41 |
TransNeXt-Base (IN-1K pretrain, DINO 1x) |
57.1 | – | – | – | – | – | – |
|
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
|
Westlake-AI/openmixup · daishiresearch/transnext · chenller/mmseg-extension
· +1 |
2023 |
| 43 |
Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale) |
57.0 | – | – | – | – | – | – |
✓ |
Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation
|
PaddlePaddle/PaddleOCR · open-mmlab/mmdetection · tensorflow/tpu
· +2 |
2020 |
| 44 |
TransNeXt-Small (IN-1K pretrain, DINO 1x) |
56.6 | – | – | – | – | – | – |
|
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
|
Westlake-AI/openmixup · daishiresearch/transnext · chenller/mmseg-extension
· +1 |
2023 |
| 45 |
QueryInst (single scale) |
56.1 | 75.8 | 61.7 | 40.2 | 59.8 | 71.5 | – |
|
Instances as Queries
|
open-mmlab/mmdetection · hustvl/QueryInst · Bo396543018/picodet_repro
· +2 |
2021 |
| 45 |
MViTv2-H (Cascade Mask R-CNN, single-scale, IN21k pre-train) |
56.1 | – | – | – | – | – | – |
|
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
|
rwightman/pytorch-image-models · facebookresearch/detectron2 · facebookresearch/SlowFast
· +6 |
2021 |
| 47 |
MOAT-0 (IN-1K pretraining, single-scale) |
55.9 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 48 |
TransNeXt-Tiny (IN-1K pretrain, DINO 1x) |
55.7 | – | – | – | – | – | – |
|
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
|
Westlake-AI/openmixup · daishiresearch/transnext · chenller/mmseg-extension
· +1 |
2023 |
| 49 |
YOLOv4-P7 CSP-P7 (single-scale, 16 fps) |
55.4 | 73.3 | 60.7 | 38.1 | 59.5 | 67.4 | – |
|
Scaled-YOLOv4: Scaling Cross Stage Partial Network
|
AlexeyAB/darknet · RangiLyu/nanodet · WongKinYiu/ScaledYOLOv4
· +38 |
2020 |
| 50 |
tiny-MOAT-3 (IN-1K pretraining, single-scale) |
55.2 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 51 |
FAN-L-Hybrid |
55.1 | – | – | – | – | – | – |
|
Understanding The Robustness in Vision Transformers
|
nvlabs/fan · NVlabs/STL |
2022 |
| 52 |
Hiera-L |
55 | – | – | – | – | – | – |
|
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
|
huggingface/pytorch-image-models · facebookresearch/hiera · leondgarse/keras_cv_attention_models
· +1 |
2023 |
| 52 |
GLEE-Lite |
55.0 | – | – | – | – | – | – |
✓ |
General Object Foundation Model for Images and Videos at Scale
|
FoundationVision/GLEE |
2023 |
| 54 |
TEC(VIT-B, Mask-RCNN) |
54.6 | – | – | – | – | – | – |
|
Towards Sustainable Self-supervised Learning
|
sail-sg/tec |
2022 |
| 55 |
Cascade Eff-B7 NAS-FPN (1280) |
54.5 | – | – | – | – | – | – |
|
Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation
|
PaddlePaddle/PaddleOCR · open-mmlab/mmdetection · tensorflow/tpu
· +2 |
2020 |
| 55 |
CAE (ViT-L, Mask R-CNN, 1x schedule) |
54.5 | – | – | – | – | – | – |
|
Context Autoencoder for Self-Supervised Representation Learning
|
open-mmlab/mmselfsup · PaddlePaddle/PaddleFL · PaddlePaddle/VIMER
· +3 |
2022 |
| 57 |
MViTv2-L (Cascade Mask R-CNN, single-scale) |
54.3 | – | – | – | – | – | – |
|
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
|
rwightman/pytorch-image-models · facebookresearch/detectron2 · facebookresearch/SlowFast
· +6 |
2021 |
| 58 |
SpineNet-190 (1280, with Self-training on OpenImages, single-scale) |
54.2 | – | – | – | – | – | – |
✓ |
Rethinking Pre-training and Self-training
|
tensorflow/tpu · stanleyjzheng/PyData |
2020 |
| 59 |
Cascade RCNN-RS (SpineNet-143L, single scale) |
53.6 | – | – | 34.5 | 56.7 | 70.6 | – |
|
Simple Training Strategies and Model Scaling for Object Detection
|
tensorflow/tpu |
2021 |
| 60 |
UniverseNet-20.08d (Res2Net-101, DCN, multi-scale) |
53.5 | 70.8 | 58.9 | 36.9 | 57.5 | 68.1 | – |
|
USB: Universal-Scale Object Detection Benchmark
|
shinya7y/UniverseNet |
2021 |
| 61 |
MAE (ViT-L, Mask R-CNN) |
53.3 | – | – | – | – | – | – |
|
Masked Autoencoders Are Scalable Vision Learners
|
facebookresearch/mae · lightly-ai/lightly · open-mmlab/mmselfsup
· +55 |
2021 |
| 62 |
Cascade RCNN-RS (ResNet-200, single scale) |
53.1 | – | – | 33.9 | 56.2 | 70.3 | – |
|
Simple Training Strategies and Model Scaling for Object Detection
|
tensorflow/tpu |
2021 |
| 63 |
tiny-MOAT-2 (IN-1K pretraining, single-scale) |
53.0 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 64 |
MViT-L (Mask R-CNN, single-scale, IN21k pre-train) |
52.7 | – | – | – | – | – | – |
|
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
|
rwightman/pytorch-image-models · facebookresearch/detectron2 · facebookresearch/SlowFast
· +6 |
2021 |
| 65 |
ResNeSt-200 (multi-scale) |
52.47 | 71.00 | 57.07 | 36.80 | 56.36 | 66.29 | – |
|
ResNeSt: Split-Attention Networks
|
rwightman/pytorch-image-models · open-mmlab/mmdetection · open-mmlab/mmpose
· +33 |
2020 |
| 66 |
ActiveMLP-B (Cascade Mask R-CNN) |
52.3 | – | – | – | – | – | – |
|
Active Token Mixer
|
microsoft/TokenMixers · microsoft/activemlp |
2022 |
| 67 |
RetinaNet (SpineNet-190, 1536x1536) |
52.2 | – | – | – | – | – | – |
|
SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
|
tensorflow/models · tensorflow/tpu · tensorflow/tpu
· +10 |
2019 |
| 68 |
EfficientDet-D7 (1536) |
52.1 | – | – | – | – | – | – |
|
EfficientDet: Scalable and Efficient Object Detection
|
tensorflow/models · PaddlePaddle/PaddleDetection · google/automl
· +61 |
2019 |
| 69 |
tiny-MOAT-1 (IN-1K pretraining, single-scale) |
51.9 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 70 |
GCNet (ResNeXt-101 + DCN + cascade + GC r4) |
51.8 | 70.4 | 56.1 | – | – | – | – |
|
Global Context Networks
|
rwightman/pytorch-image-models · PaddlePaddle/PaddleDetection · xvjiarui/GCNet |
2020 |
| 71 |
ELSA-S (Cascade Mask RCNN) |
51.6 | 70.5 | 56.0 | – | – | – | – |
|
ELSA: Enhanced Local Self-Attention for Vision Transformer
|
damo-cv/elsa |
2021 |
| 72 |
FocalNet-T (LRF, Cascade Mask R-CNN) |
51.5 | 70.3 | 56.0 | – | – | – | – |
|
Focal Modulation Networks
|
PaddlePaddle/PaddleDetection · keras-team/keras-io · microsoft/FocalNet
· +6 |
2022 |
| 73 |
DINO-5scale (24 epoch) |
51.3 | 69.1 | 56 | 34.5 | 54.2 | 65.8 | – |
|
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
|
IDEA-Research/Grounded-Segment-Anything · PaddlePaddle/PaddleDetection · lucasjinreal/yolov7_d2
· +13 |
2022 |
| 74 |
DINO-5scale (36 epoch) |
51.2 | 69 | 55.8 | 35 | 54.3 | 65.3 | – |
|
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
|
IDEA-Research/Grounded-Segment-Anything · PaddlePaddle/PaddleDetection · lucasjinreal/yolov7_d2
· +13 |
2022 |
| 75 |
ResNeSt-200-DCN (single-scale) |
50.91 | 69.53 | 55.40 | 32.67 | 54.66 | 65.83 | – |
|
ResNeSt: Split-Attention Networks
|
rwightman/pytorch-image-models · open-mmlab/mmdetection · open-mmlab/mmpose
· +33 |
2020 |
| 76 |
UniverseNet-20.08d (Res2Net-101, DCN, single-scale) |
50.9 | 69.5 | 55.4 | 33.5 | 55.5 | 65.8 | – |
|
USB: Universal-Scale Object Detection Benchmark
|
shinya7y/UniverseNet |
2021 |
| 77 |
ResNeSt-200 (single-scale) |
50.54 | 68.78 | 55.17 | – | 54.2 | 63.9 | – |
|
ResNeSt: Split-Attention Networks
|
rwightman/pytorch-image-models · open-mmlab/mmdetection · open-mmlab/mmpose
· +33 |
2020 |
| 78 |
tiny-MOAT-0 (IN-1K pretraining, single-scale) |
50.5 | – | – | – | – | – | – |
|
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
|
google-research/deeplab2 · RooKichenn/pytorch-MOAT |
2022 |
| 79 |
MAE (ViT-B, Mask R-CNN) |
50.3 | – | – | – | – | – | – |
|
Masked Autoencoders Are Scalable Vision Learners
|
facebookresearch/mae · lightly-ai/lightly · open-mmlab/mmselfsup
· +55 |
2021 |
| 80 |
Sparse R-CNN (PVTv2-B2) |
50.1 | 69.5 | 54.9 | – | – | – | – |
|
PVT v2: Improved Baselines with Pyramid Vision Transformer
|
rwightman/pytorch-image-models · open-mmlab/mmdetection · open-mmlab/mmpose
· +15 |
2021 |
| 81 |
Pix2seq (ViT-L) |
50.0 | – | – | – | – | – | – |
✓ |
Pix2seq: A Language Modeling Framework for Object Detection
|
google-research/pix2seq · gaopengcuhk/Stable-Pix2Seq · gaopengcuhk/Unofficial-Pix2Seq
· +3 |
2021 |
| 82 |
DaViT-T (Mask R-CNN, 36 epochs) |
49.9 | – | – | – | – | – | – |
|
DaViT: Dual Attention Vision Transformers
|
rwightman/pytorch-image-models · leondgarse/keras_cv_attention_models · dingmyu/davit
· +1 |
2022 |
| 83 |
BoTNet 200 (Mask R-CNN, single scale, 72 epochs) |
49.7 | 71.3 | 54.6 | – | – | – | – |
|
Bottleneck Transformers for Visual Recognition
|
rwightman/pytorch-image-models · BR-IDL/PaddleViT · The-AI-Summer/self_attention
· +10 |
2021 |
| 84 |
BoTNet 152 (Mask R-CNN, single scale, 72 epochs) |
49.5 | 71 | 54.2 | – | – | – | – |
|
Bottleneck Transformers for Visual Recognition
|
rwightman/pytorch-image-models · BR-IDL/PaddleViT · The-AI-Summer/self_attention
· +10 |
2021 |
| 84 |
DN-Deformable-DETR-R50++ |
49.5 | 67.6 | 53.8 | 31.3 | 52.6 | 65.4 | 47 |
|
DN-DETR: Accelerate DETR Training by Introducing Query DeNoising
|
IDEACVR/DINO · idea-research/dino · IDEA-Research/detrex
· +14 |
2022 |
| 86 |
REGO-Deformable DETR-X101 |
49.1 | 67.5 | 53.1 | 30 | 52.6 | 65 | – |
|
Recurrent Glimpse-based Decoder for Detection with Transformer
|
zhechen/deformable-detr-rego |
2021 |
| 87 |
CenterMask+VoVNet99 (multi-scale) |
48.6 | 67.8 | – | – | – | – | – |
|
CenterMask : Real-Time Anchor-Free Instance Segmentation
|
youngwanLEE/centermask2 · youngwanLEE/CenterMask · youngwanLEE/vovnet-detectron2
· +5 |
2019 |
| 87 |
Mask R-CNN (ResNeXt-152-FPN, cascade) |
48.6 | 66.8 | 52.9 | – | – | – | – |
|
Rethinking ImageNet Pre-training
|
tensorpack/tensorpack |
2018 |
| 89 |
UniverseNet-20.08 (Res2Net-50, DCN, single-scale) |
48.5 | 67.0 | 52.6 | 30.6 | 52.7 | 62.7 | – |
|
USB: Universal-Scale Object Detection Benchmark
|
shinya7y/UniverseNet |
2021 |
| 89 |
XCiT-M24/8 |
48.5 | – | – | – | – | – | – |
|
XCiT: Cross-Covariance Image Transformers
|
rwightman/pytorch-image-models · facebookresearch/dino · facebookresearch/vissl
· +9 |
2021 |
| 91 |
ELSA-S (Mask RCNN) |
48.3 | 70.4 | 52.9 | – | – | – | – |
|
ELSA: Enhanced Local Self-Attention for Vision Transformer
|
damo-cv/elsa |
2021 |
| 92 |
XCiT-S24/8 |
48.1 | – | – | – | – | – | – |
|
XCiT: Cross-Covariance Image Transformers
|
rwightman/pytorch-image-models · facebookresearch/dino · facebookresearch/vissl
· +9 |
2021 |
| 93 |
GCNet (ResNeXt-101 + DCN + cascade + GC r16) |
47.9 | 66.9 | 52.2 | – | – | – | – |
|
GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
|
open-mmlab/mmdetection · open-mmlab/mmsegmentation · PaddlePaddle/PaddleSeg
· +6 |
2019 |
| 94 |
MAE-Det(MAE-Det-L+GFLV2) |
47.8 | 65.5 | 52.2 | 30.3 | 51.9 | 61.1 | – |
|
MAE-DET: Revisiting Maximum Entropy Principle in Zero-Shot NAS for Efficient Object Detection
|
alibaba/lightweight-neural-architecture-search |
2021 |
| 95 |
Res2Net101+HTC |
47.5 | 66.5 | 51.3 | 28.6 | 51.6 | 62.1 | – |
|
Res2Net: A New Multi-scale Backbone Architecture
|
rwightman/pytorch-image-models · open-mmlab/mmdetection · PaddlePaddle/PaddleDetection
· +31 |
2019 |
| 96 |
Mask R-CNN (ResNet-101-FPN, GN, Cascade) |
47.4 | – | – | – | – | – | – |
|
Rethinking ImageNet Pre-training
|
tensorpack/tensorpack |
2018 |
| 97 |
Pix2seq (R50-C4) |
47.3 | – | – | – | – | – | – |
|
Pix2seq: A Language Modeling Framework for Object Detection
|
google-research/pix2seq · gaopengcuhk/Stable-Pix2Seq · gaopengcuhk/Unofficial-Pix2Seq
· +3 |
2021 |
| 98 |
Pix2seq (ViT-B) |
47.1 | – | – | – | – | – | – |
|
Pix2seq: A Language Modeling Framework for Object Detection
|
google-research/pix2seq · gaopengcuhk/Stable-Pix2Seq · gaopengcuhk/Unofficial-Pix2Seq
· +3 |
2021 |
| 99 |
HTC (HRNetV2p-W48) |
47.0 | – | – | 28.8 | 50.3 | 62.2 | – |
|
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 99 |
PatchConvNet-S120 (Mask R-CNN) |
47.0 | – | – | – | – | – | – |
|
Augmenting Convolutional networks with attention-based aggregation
|
facebookresearch/deit · keras-team/keras-io · DarshanDeshpande/jax-models
· +2 |
2021 |