| Rank | Model |
box AP | AP50 | AP75 | APS | APM | APL | Params (M) |
Paper | Code | Year |
| 101 |
RPDet (ResNeXt-101-DCN, multi-scale) |
46.8 | – | – | – | – | – | – |
RepPoints: Point Set Representation for Object Detection
|
open-mmlab/mmdetection · microsoft/RepPoints · Scalsol/RepPointsV2
· +3 |
2019 |
| 102 |
DAB-DETR-DC5-R101 |
46.6 | 67 | 50.2 | 28.1 | 50.5 | 64.1 | 63 |
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
|
IDEA-Research/detrex · alibaba/EasyCV · idea-research/dn-detr
· +5 |
2022 |
| 103 |
DyHead (ResNet-101) |
46.5 | – | – | – | – | – | – |
Dynamic Head: Unifying Object Detection Heads with Attentions
|
open-mmlab/mmdetection · microsoft/DynamicHead · Coldestadam/DynamicHead |
2021 |
| 104 |
Mask R-CNN (ResNeXt-152-FPN) |
46.4 | 67.1 | 51.1 | – | – | – | – |
Rethinking ImageNet Pre-training
|
tensorpack/tensorpack |
2018 |
| 104 |
RPDet (ResNet-101-DCN, multi-scale) |
46.4 | – | – | – | – | – | – |
RepPoints: Point Set Representation for Object Detection
|
open-mmlab/mmdetection · microsoft/RepPoints · Scalsol/RepPointsV2
· +3 |
2019 |
| 104 |
PatchConvNet-S60 (Mask R-CNN) |
46.4 | – | – | – | – | – | – |
Augmenting Convolutional networks with attention-based aggregation
|
facebookresearch/deit · keras-team/keras-io · DarshanDeshpande/jax-models
· +2 |
2021 |
| 107 |
Cascade Mask R-CNN (ResNet-50) |
46.3 | 64.3 | 50.5 | – | – | – | – |
Deep Residual Learning for Image Recognition
|
tensorflow/models · tensorflow/models · tensorflow/models
· +481 |
2015 |
| 108 |
HoughNet (HG-104, MS) |
46.1 | 64.6 | 50.3 | 30.0 | 48.8 | 59.7 | – |
HoughNet: Integrating near and long-range evidence for bottom-up object detection
|
giddyyupp/coco-minitrain · nerminsamet/houghnet |
2020 |
| 109 |
Mask R-CNN (HRNetV2p-W48, cascade) |
46.0 | – | – | 27.5 | – | 60.1 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 110 |
Conditional DETR-DC5-R101 |
45.9 | 66.8 | 49.5 | 27.2 | 50.3 | 63.3 | 63 |
Conditional DETR for Fast Training Convergence
|
huggingface/transformers · IDEA-Research/detrex · atten4vis/conditionaldetr
· +1 |
2021 |
| 110 |
BoTNet 50 (72 epochs) |
45.9 | – | – | – | – | – | – |
Bottleneck Transformers for Visual Recognition
|
rwightman/pytorch-image-models · BR-IDL/PaddleViT · The-AI-Summer/self_attention
· +10 |
2021 |
| 112 |
Sparse R-CNN (ResNet-101, learnable proposals, random crop aug, FPN) |
45.6 | 64.6 | 49.5 | 28.3 | 48.3 | 61.6 | – |
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · PeizeSun/SparseR-CNN
· +3 |
2020 |
| 112 |
CenterMask+VoVNetV2-99 (single-scale) |
45.6 | – | – | 29.2 | – | 58.8 | – |
CenterMask : Real-Time Anchor-Free Instance Segmentation
|
youngwanLEE/centermask2 · youngwanLEE/CenterMask · youngwanLEE/vovnet-detectron2
· +5 |
2019 |
| 114 |
HTC (HRNetV2p-W32) |
45.3 | – | – | 27.0 | 48.4 | 59.5 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 115 |
Anchor DETR-DC5-R101 |
45.1 | 65.7 | 48.8 | 25.8 | 49.4 | 61.6 | – |
Anchor DETR: Query Design for Transformer-Based Object Detection
|
megvii-research/AnchorDETR · megvii-model/anchordetr |
2021 |
| 115 |
Conditional DETR-DC5-R50 |
45.1 | 65.4 | 48.5 | 25.3 | 49 | 62.2 | 44 |
Conditional DETR for Fast Training Convergence
|
huggingface/transformers · IDEA-Research/detrex · atten4vis/conditionaldetr
· +1 |
2021 |
| 117 |
Mask R-CNN (ResNeXt-152 + 1 NL) |
45.0 | 67.8 | 48.9 | – | – | – | – |
Non-local Neural Networks
|
facebookresearch/detectron · facebookresearch/SlowFast · open-mmlab/mmaction2
· +29 |
2017 |
| 117 |
Pix2seq (R101-DC5) |
45.0 | 63.2 | 48.6 | 28.2 | 48.9 | 60.4 | – |
Pix2seq: A Language Modeling Framework for Object Detection
|
google-research/pix2seq · gaopengcuhk/Stable-Pix2Seq · gaopengcuhk/Unofficial-Pix2Seq
· +3 |
2021 |
| 119 |
Mask R-CNN-FPN (AOGNet-40M) |
44.9 | 66.2 | 49.1 | – | – | – | – |
Attentive Normalization
|
iVMCL/AOGNet-v2 · ivMCL/AttentiveNorm_Detection |
2019 |
| 119 |
DETR-DC5 (ResNet-101) |
44.9 | 64.7 | 47.7 | 23.7 | 49.5 | 62.3 | – |
End-to-End Object Detection with Transformers
|
huggingface/transformers · tensorflow/models · open-mmlab/mmdetection
· +34 |
2020 |
| 119 |
Mask R-CNN (VoVNetV2-99, single-scale) |
44.9 | – | – | 28.5 | – | 57.7 | – |
CenterMask : Real-Time Anchor-Free Instance Segmentation
|
youngwanLEE/centermask2 · youngwanLEE/CenterMask · youngwanLEE/vovnet-detectron2
· +5 |
2019 |
| 122 |
R3-CNN (ResNet-50-FPN, DCN) |
44.8 | 64.3 | 48.9 | 26.6 | 48.3 | 59.6 | – |
Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing
|
IMPLabUniPr/mmdetection |
2021 |
| 122 |
RPDet (ResNet-101-DCN, multi-scale train) |
44.8 | – | – | – | – | – | – |
RepPoints: Point Set Representation for Object Detection
|
open-mmlab/mmdetection · microsoft/RepPoints · Scalsol/RepPointsV2
· +3 |
2019 |
| 124 |
RetinaNet (ViL-Base, multi-scale, 3x) |
44.7 | – | 47.6 | 29.9 | 48 | 58.1 | – |
Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
|
microsoft/esvit · microsoft/vision-longformer · microsoft/VisionLongformerForObjectDetection |
2021 |
| 125 |
Cascade R-CNN (HRNetV2p-W48) |
44.6 | 62.7 | 48.7 | 26.3 | 48.1 | 58.5 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 125 |
CenterMask+VoVNetV2-57 (single-scale) |
44.6 | – | – | 27.7 | 48.3 | – | – |
CenterMask : Real-Time Anchor-Free Instance Segmentation
|
youngwanLEE/centermask2 · youngwanLEE/CenterMask · youngwanLEE/vovnet-detectron2
· +5 |
2019 |
| 127 |
Conditional DETR-R101 |
44.5 | 65.6 | 47.5 | 23.6 | 48.4 | 63.6 | 63 |
Conditional DETR for Fast Training Convergence
|
huggingface/transformers · IDEA-Research/detrex · atten4vis/conditionaldetr
· +1 |
2021 |
| 127 |
Sparse R-CNN (ResNet-50, learnable proposals, random crop aug, FPN) |
44.5 | 63.4 | 48.2 | 26.9 | 47.2 | 59.5 | – |
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · PeizeSun/SparseR-CNN
· +3 |
2020 |
| 127 |
GFL (ResNet-50) |
44.5 | 63.0 | 48.3 | – | – | – | – |
Deep Residual Learning for Image Recognition
|
tensorflow/models · tensorflow/models · tensorflow/models
· +481 |
2015 |
| 127 |
RPDet (ResNeXt-101-DCN) |
44.5 | – | – | – | – | – | – |
RepPoints: Point Set Representation for Object Detection
|
open-mmlab/mmdetection · microsoft/RepPoints · Scalsol/RepPointsV2
· +3 |
2019 |
| 131 |
CenterMask+X101-32x8d (single-scale) |
44.4 | – | – | 26.7 | – | 57.1 | – |
CenterMask : Real-Time Anchor-Free Instance Segmentation
|
youngwanLEE/centermask2 · youngwanLEE/CenterMask · youngwanLEE/vovnet-detectron2
· +5 |
2019 |
| 132 |
RetinaNet (ViL-Base) |
44.3 | 65.5 | 47.1 | 28.9 | 47.9 | 58.3 | – |
Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
|
microsoft/esvit · microsoft/vision-longformer · microsoft/VisionLongformerForObjectDetection |
2021 |
| 132 |
R3-CNN (ResNet-50-FPN, GC-Net) |
44.3 | 64.1 | 48.4 | 27 | 47.1 | 58.9 | – |
Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing
|
IMPLabUniPr/mmdetection |
2021 |
| 134 |
Anchor DETR-DC5-R50 |
44.2 | 64.7 | 47.5 | 24.7 | 48.2 | 60.6 | – |
Anchor DETR: Query Design for Transformer-Based Object Detection
|
megvii-research/AnchorDETR · megvii-model/anchordetr |
2021 |
| 135 |
DAB-DETR-R101 |
44.1 | 64.7 | 47.2 | 24.1 | 48.2 | 62.9 | 63 |
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
|
IDEA-Research/detrex · alibaba/EasyCV · idea-research/dn-detr
· +5 |
2022 |
| 136 |
Faster RCNN-R101-FPN+ |
44 | 63.9 | 47.8 | 27.2 | 48.1 | 56 | – |
End-to-End Object Detection with Transformers
|
huggingface/transformers · tensorflow/models · open-mmlab/mmdetection
· +34 |
2020 |
| 137 |
Cascade R-CNN (HRNetV2p-W32) |
43.7 | 61.7 | 47.7 | 25.6 | 46.5 | 57.4 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 138 |
Sparse R-CNN (ResNet-101, FPN) |
43.5 | 62.1 | 47.2 | 26.1 | 46.3 | 59.7 | – |
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · PeizeSun/SparseR-CNN
· +3 |
2020 |
| 138 |
ATSS (ResNet-50) |
43.5 | 61.9 | 47.0 | – | – | – | – |
Deep Residual Learning for Image Recognition
|
tensorflow/models · tensorflow/models · tensorflow/models
· +481 |
2015 |
| 140 |
PVT-Large (RetinaNet 3x,MS) |
43.4 | 63.6 | 46.1 | 26.1 | 46.0 | 59.5 | – |
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
|
open-mmlab/mmdetection · open-mmlab/mmpose · whai362/PVT
· +8 |
2021 |
| 141 |
ExtremeNet (Hourglass-104, multi-scale) |
43.3 | 59.6 | 46.8 | 25.7 | 46.6 | 59.4 | – |
Bottom-up Object Detection by Grouping Extreme and Center Points
|
xingyizhou/ExtremeNet · DataXujing/ExtremeNet-Pytorch |
2019 |
| 142 |
Pix2seq (R50-DC5 ) |
43.2 | 61.0 | 46.1 | 26.6 | 47 | 58.6 | – |
Pix2seq: A Language Modeling Framework for Object Detection
|
google-research/pix2seq · gaopengcuhk/Stable-Pix2Seq · gaopengcuhk/Unofficial-Pix2Seq
· +3 |
2021 |
| 142 |
HTC (cascade) |
43.2 | 59.4 | 40.7 | 20.3 | 40.9 | 52.3 | – |
Hybrid Task Cascade for Instance Segmentation
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · amirassov/kaggle-imaterialist
· +2 |
2019 |
| 144 |
Mask R-CNN-FPN (ResNeXt-101, GN+WS) |
43.12 | 64.15 | 47.11 | 25.49 | 47.19 | 56.39 | – |
Micro-Batch Training with Batch-Channel Normalization and Weight Standardization
|
labmlai/annotated_deep_learning_paper_implementations · joe-siyuan-qiao/WeightStandardization · jinfagang/nb
· +4 |
2019 |
| 145 |
HTC (HRNetV2p-W18) |
43.1 | – | – | 26.6 | 46.0 | – | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 145 |
Mask R-CNN (ResNet-101, DCNv2) |
43.1 | – | – | – | – | – | – |
Deformable ConvNets v2: More Deformable, Better Results
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · msracver/Deformable-ConvNets
· +23 |
2018 |
| 147 |
Conditional DETR-R50 |
43 | 64 | 45.7 | 22.7 | 46.7 | 61.5 | 44 |
Conditional DETR for Fast Training Convergence
|
huggingface/transformers · IDEA-Research/detrex · atten4vis/conditionaldetr
· +1 |
2021 |
| 147 |
HoughNet (HG-104) |
43.0 | 62.2 | 46.9 | 25.5 | 47.6 | 55.8 | – |
HoughNet: Integrating near and long-range evidence for bottom-up object detection
|
giddyyupp/coco-minitrain · nerminsamet/houghnet |
2020 |
| 149 |
Faster R-CNN (FPN, X-volution) |
42.8 | 64 | 46.4 | 26.9 | 46 | 55 | – |
X-volution: On the unification of convolution and self-attention
|
|
2021 |
| 150 |
Cascade R-CNN (ResNet-101-FPN+, cascade) |
42.7 | 61.6 | 46.6 | 23.8 | 46.2 | 57.4 | – |
Cascade R-CNN: Delving into High Quality Object Detection
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · tensorpack/tensorpack
· +5 |
2017 |
| 151 |
PVT-Large (RetinaNet 1x) |
42.6 | 63.7 | 45.4 | 25.8 | 46.0 | 58.4 | – |
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
|
open-mmlab/mmdetection · open-mmlab/mmpose · whai362/PVT
· +8 |
2021 |
| 151 |
CornerNet-Saccade (Hourglass-54) |
42.6 | – | – | 25.5 | 44.3 | 58.4 | – |
CornerNet-Lite: Efficient Keypoint Based Object Detection
|
PaddlePaddle/PaddleDetection · princeton-vl/CornerNet-Lite · takooctopus/CornerNet-Lite-Tako
· +3 |
2019 |
| 151 |
Pix2seq (R50) |
42.6 | – | – | – | – | – | – |
Pix2seq: A Language Modeling Framework for Object Detection
|
google-research/pix2seq · gaopengcuhk/Stable-Pix2Seq · gaopengcuhk/Unofficial-Pix2Seq
· +3 |
2021 |
| 154 |
Mask R-CNN (ResNet-101-FPN, GroupNorm, long) |
42.3 | 62.8 | 46.2 | – | – | – | – |
Group Normalization
|
labmlai/annotated_deep_learning_paper_implementations · facebookresearch/detectron · PaddlePaddle/PaddleDetection
· +19 |
2018 |
| 154 |
Sparse R-CNN (ResNet-50, FPN) |
42.3 | 61.2 | 45.7 | 26.7 | 44.6 | 57.6 | – |
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · PeizeSun/SparseR-CNN
· +3 |
2020 |
| 154 |
Mask R-CNN (HRNetV2p-W32) |
42.3 | – | – | 25.0 | 45.4 | – | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 154 |
DETR-ResNet50 with iRPE-K (300 epochs) |
42.3 | – | – | – | – | – | – |
Rethinking and Improving Relative Position Encoding for Vision Transformer
|
microsoft/cream |
2021 |
| 158 |
TridentNet (ResNet-101) |
42 | 63.5 | 45.5 | 24.9 | 47 | 56.9 | – |
Scale-Aware Trident Networks for Object Detection
|
facebookresearch/detectron2 · open-mmlab/mmdetection · tusimple/simpledet
· +1 |
2019 |
| 158 |
R3-CNN (ResNet-50-FPN) |
42 | 61 | 46.3 | 24.5 | 45.2 | 55.7 | – |
Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing
|
IMPLabUniPr/mmdetection |
2021 |
| 160 |
Faster R-CNN (HRNetV2p-W48) |
41.8 | 62.8 | 45.9 | – | 44.7 | 54.6 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 161 |
Faster R-CNN (LIP-ResNet-101) |
41.7 | 63.6 | 45.6 | 25.2 | 45.8 | – | – |
LIP: Local Importance-based Pooling
|
sebgao/LIP |
2019 |
| 161 |
Faster R-CNN (ResNet-101, DCNv2) |
41.7 | – | – | 22.2 | 45.8 | 58.7 | – |
Deformable ConvNets v2: More Deformable, Better Results
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · msracver/Deformable-ConvNets
· +23 |
2018 |
| 163 |
FSAF (ResNeXt-101, anchor-based branches) |
41.6 | 62.4 | – | – | – | – | – |
Feature Selective Anchor-Free Module for Single-Shot Object Detection
|
open-mmlab/mmdetection · hdjang/Feature-Selective-Anchor-Free-Module-for-Single-Shot-Object-Detection · xuannianz/FSAF
· +1 |
2019 |
| 164 |
CornerNet-Saccade (Hourglass-104) |
41.4 | – | – | 23.8 | 43.5 | 57.1 | – |
CornerNet-Lite: Efficient Keypoint Based Object Detection
|
PaddlePaddle/PaddleDetection · princeton-vl/CornerNet-Lite · takooctopus/CornerNet-Lite-Tako
· +3 |
2019 |
| 165 |
Grid R-CNN (ResNet-101-FPN) |
41.3 | 60.3 | 44.4 | 23.4 | 45.8 | 54.1 | – |
Grid R-CNN
|
open-mmlab/mmdetection · STVIR/Grid-R-CNN |
2018 |
| 165 |
Cascade R-CNN (HRNetV2p-W18) |
41.3 | 59.2 | 44.9 | 23.7 | 44.2 | 54.1 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 165 |
CenterNet511 (Hourglass-52) |
41.3 | 59.2 | 43.9 | 23.6 | 43.8 | 55.8 | – |
CenterNet: Keypoint Triplets for Object Detection
|
Duankaiwen/CenterNet · ximilar-com/xcenternet · kuku-sichuan/CenterNet
· +17 |
2019 |
| 168 |
RetinaMask (ResNet-101-FPN) |
41.1 | 60.2 | 44.1 | – | – | – | – |
RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
|
chengyangfu/retinamask · lzrobots/dgmn · oulutan/Drone_FasterRCNN
· +50 |
2019 |
| 169 |
PoolFormer-S36 (Mask R-CNN) |
41.0 | 63.1 | 44.8 | – | – | – | – |
MetaFormer Is Actually What You Need for Vision
|
huggingface/transformers · rwightman/pytorch-image-models · facebookresearch/xformers
· +15 |
2021 |
| 170 |
Faster R-CNN (HRNetV2p-W32) |
40.9 | 61.8 | 44.8 | 24.4 | 43.7 | 53.3 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 170 |
VirTex Mask R-CNN (ResNet-50-FPN) |
40.9 | – | – | – | – | – | – |
VirTex: Learning Visual Representations from Textual Annotations
|
kdexd/virtex · mattdeitke/cvpr-buzz · rahulvigneswaran/longtail-buzz |
2020 |
| 172 |
Mask R-CNN (ResNet-101 + 1 NL) |
40.8 | 63.1 | 44.5 | – | – | – | – |
Non-local Neural Networks
|
facebookresearch/detectron · facebookresearch/SlowFast · open-mmlab/mmaction2
· +29 |
2017 |
| 172 |
Mask R-CNN (ResNet-50-FPN, GroupNorm, long) |
40.8 | 61.6 | 44.4 | – | – | – | – |
Group Normalization
|
labmlai/annotated_deep_learning_paper_implementations · facebookresearch/detectron · PaddlePaddle/PaddleDetection
· +19 |
2018 |
| 172 |
RPDet (ResNet-50, multi-scale train) |
40.8 | – | – | – | – | – | – |
RepPoints: Point Set Representation for Object Detection
|
open-mmlab/mmdetection · microsoft/RepPoints · Scalsol/RepPointsV2
· +3 |
2019 |
| 172 |
DETR-ResNet50 with iRPE-K (150 epochs) |
40.8 | – | – | – | – | – | – |
Rethinking and Improving Relative Position Encoding for Vision Transformer
|
microsoft/cream |
2021 |
| 176 |
Faster R-CNN+aLRP Loss (ResNet-50, 500 scale) |
40.7 | 60.7 | 43.3 | – | – | – | – |
A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection
|
kemaloksuz/aLRPLoss · xudangliatiger/ape-loss · kemaloksuz/aLRPLoss-AblationExperiments |
2020 |
| 177 |
PPDet (ResNet-101-FPN) |
40.5 | 59.5 | 44.2 | 25.4 | 44.7 | 52.3 | – |
Reducing Label Noise in Anchor-Free Object Detection
|
nerminsamet/ppdet |
2020 |
| 178 |
GCnet (ResNet-50-FPN, GRoIE) |
40.3 | 62.4 | 44 | 24.2 | 44.4 | 52.5 | – |
GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
|
open-mmlab/mmdetection · open-mmlab/mmsegmentation · PaddlePaddle/PaddleSeg
· +6 |
2019 |
| 178 |
Mask R-CNN (ResNet-50-FPN, GroupNorm) |
40.3 | 61 | 44 | – | – | – | – |
Group Normalization
|
labmlai/annotated_deep_learning_paper_implementations · facebookresearch/detectron · PaddlePaddle/PaddleDetection
· +19 |
2018 |
| 178 |
Cascade R-CNN (ResNet-50-FPN+) |
40.3 | 59.4 | 43.7 | 22.9 | 43.7 | 54.1 | – |
Cascade R-CNN: Delving into High Quality Object Detection
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · tensorpack/tensorpack
· +5 |
2017 |
| 178 |
ExtremeNet (Hourglass-104, single-scale) |
40.3 | 55.1 | 43.7 | 21.6 | 44.0 | 56.1 | – |
Bottom-up Object Detection by Grouping Extreme and Center Points
|
xingyizhou/ExtremeNet · DataXujing/ExtremeNet-Pytorch |
2019 |
| 178 |
RPDet (ResNet-101) |
40.3 | – | – | – | – | – | – |
RepPoints: Point Set Representation for Object Detection
|
open-mmlab/mmdetection · microsoft/RepPoints · Scalsol/RepPointsV2
· +3 |
2019 |
| 183 |
RetinaNet+aLRP Loss (ResNet-50, 500 scale) |
40.2 | 60.3 | 42.3 | – | – | – | – |
A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection
|
kemaloksuz/aLRPLoss · xudangliatiger/ape-loss · kemaloksuz/aLRPLoss-AblationExperiments |
2020 |
| 184 |
Mask R-CNN (ResNet-101-FPN) |
40.0 | – | – | – | – | – | – |
Mask R-CNN
|
tensorflow/models · facebookresearch/detectron2 · facebookresearch/detectron
· +176 |
2017 |
| 185 |
FPN+ |
39.8 | 61.3 | 43.3 | 22.9 | 43.3 | 52.6 | – |
Feature Pyramid Networks for Object Detection
|
PaddlePaddle/PaddleOCR · open-mmlab/mmdetection · facebookresearch/detectron
· +82 |
2016 |
| 186 |
FoveaBox+aLRP Loss (ResNet-50, 500 scale) |
39.7 | 58.8 | 41.5 | – | – | – | – |
A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection
|
kemaloksuz/aLRPLoss · xudangliatiger/ape-loss · kemaloksuz/aLRPLoss-AblationExperiments |
2020 |
| 187 |
Grid R-CNN (ResNet-50-FPN) |
39.6 | 58.3 | 42.4 | 22.6 | 43.8 | 51.5 | – |
Grid R-CNN
|
open-mmlab/mmdetection · STVIR/Grid-R-CNN |
2018 |
| 188 |
Mask R-CNN (ResNet-50, ACNet) |
39.5 | – | – | – | – | – | – |
Adaptively Connected Neural Networks
|
wanggrun/Adaptively-Connected-Neural-Networks |
2019 |
| 189 |
FSAF (ResNet-101, anchor-based branches) |
39.3 | 59.2 | – | – | – | – | – |
Feature Selective Anchor-Free Module for Single-Shot Object Detection
|
open-mmlab/mmdetection · hdjang/Feature-Selective-Anchor-Free-Module-for-Single-Shot-Object-Detection · xuannianz/FSAF
· +1 |
2019 |
| 190 |
Mask R-CNN (HRNetV2p-W18) |
39.2 | – | – | – | 41.7 | 51.0 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 191 |
Mask R-CNN (ResNet-50 + 1 NL) |
39.0 | 61.1 | 41.9 | – | – | – | – |
Non-local Neural Networks
|
facebookresearch/detectron · facebookresearch/SlowFast · open-mmlab/mmaction2
· +29 |
2017 |
| 192 |
FoveaBox (ResNet-101-FPN, 800x800) |
38.9 | 58.4 | 41.5 | 22.3 | 43.5 | 51.7 | – |
FoveaBox: Beyond Anchor-based Object Detector
|
open-mmlab/mmdetection · taokong/FoveaBox · anonymous2020new/iffDetector
· +4 |
2019 |
| 193 |
FCOS (ResNet-50-FPN + improvements) |
38.6 | 57.4 | 41.4 | 22.3 | 42.5 | 49.8 | – |
FCOS: Fully Convolutional One-Stage Object Detection
|
open-mmlab/mmdetection · pytorch/vision · PaddlePaddle/PaddleDetection
· +84 |
2019 |
| 193 |
RPDet (ResNet-50) |
38.6 | – | – | – | – | – | – |
RepPoints: Point Set Representation for Object Detection
|
open-mmlab/mmdetection · microsoft/RepPoints · Scalsol/RepPointsV2
· +3 |
2019 |
| 195 |
Libra R-CNN (ResNet-50 FPN) |
38.5 | 59.3 | 42.0 | 22.9 | 42.1 | 50.5 | – |
Libra R-CNN: Towards Balanced Learning for Object Detection
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · OceanPang/Libra_R-CNN
· +3 |
2019 |
| 196 |
Mask R-CNN (ResNet-50-FPN, GRoIE) |
38.4 | 59.9 | 41.7 | 22.9 | 42.1 | 49.7 | – |
A novel Region of Interest Extraction Layer for Instance Segmentation
|
open-mmlab/mmdetection · open-mmlab/mmdetection · IMPLabUniPr/mmdetection-groie
· +2 |
2020 |
| 196 |
CornerNet511 (Hourglass-104) |
38.4 | 53.8 | 40.9 | 18.6 | 40.5 | 51.8 | – |
CornerNet: Detecting Objects as Paired Keypoints
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · princeton-vl/CornerNet
· +2 |
2018 |
| 198 |
FoveaBox+Retina (ResNet-50) |
38.1 | 57.8 | 40.5 | – | – | – | – |
FoveaBox: Beyond Anchor-based Object Detector
|
open-mmlab/mmdetection · taokong/FoveaBox · anonymous2020new/iffDetector
· +4 |
2019 |
| 199 |
Faster R-CNN (HRNetV2p-W18) |
38.0 | 58.9 | 41.5 | 22.6 | 40.8 | 49.6 | – |
Deep High-Resolution Representation Learning for Visual Recognition
|
open-mmlab/mmdetection · PaddlePaddle/PaddleDetection · open-mmlab/mmsegmentation
· +39 |
2019 |
| 199 |
FoveaBox (ResNet-101-FPN, 600x600) |
38 | 57.8 | 40.2 | 19.5 | 42.2 | 52.7 | – |
FoveaBox: Beyond Anchor-based Object Detector
|
open-mmlab/mmdetection · taokong/FoveaBox · anonymous2020new/iffDetector
· +4 |
2019 |