| Rank | Model |
AP | mAP | Mean mAP | AP50 | AP75 | APM | APM50 | APM75 |
Paper | Code | Year |
| 1 |
MaxViT-B |
53.4 | – | – | 72.9 | 58.1 | 45.7 | 70.3 | 50 |
MaxViT: Multi-Axis Vision Transformer
|
huggingface/pytorch-image-models · lucidrains/vit-pytorch · lucidrains/imagen-pytorch
· +12 |
2022 |
| 2 |
MaxViT-S |
53.1 | – | – | 72.5 | 58.1 | 45.4 | 69.8 | 49.5 |
MaxViT: Multi-Axis Vision Transformer
|
huggingface/pytorch-image-models · lucidrains/vit-pytorch · lucidrains/imagen-pytorch
· +12 |
2022 |
| 3 |
MaxViT-T |
52.1 | – | – | 71.9 | 56.8 | 44.6 | 69.1 | 48.4 |
MaxViT: Multi-Axis Vision Transformer
|
huggingface/pytorch-image-models · lucidrains/vit-pytorch · lucidrains/imagen-pytorch
· +12 |
2022 |
| 4 |
DAT-S++ |
50.2 | – | – | – | – | – | – | – |
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention
|
leaplabthu/dat |
2023 |
| 5 |
DAT-T++ |
49.2 | – | – | – | – | – | – | – |
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention
|
leaplabthu/dat |
2023 |
| 6 |
DyHead (SAP) |
42.1 | – | – | 59.4 | 45.9 | – | – | – |
Stochastic Subsampling With Average Pooling
|
|
2024 |
| 7 |
Faster R-CNN (ideal number of groups) |
40.7 | – | – | 61.2 | 44.6 | – | – | – |
On the Ideal Number of Groups for Isometric Gradient Propagation
|
|
2023 |
| 8 |
UniRepLKNet-XL++ |
– | 56.4 | – | – | – | – | – | – |
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
|
ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension |
2023 |
| 9 |
UniRepLKNet-L++ |
– | 55.8 | – | – | – | – | – | – |
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
|
ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension |
2023 |
| 10 |
UniRepLKNet-B++ |
– | 54.8 | – | – | – | – | – | – |
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
|
ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension |
2023 |
| 11 |
UniRepLKNet-S++ |
– | 54.3 | – | – | – | – | – | – |
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
|
ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension |
2023 |
| 12 |
MixMIM-L |
– | 54.1 | – | – | – | – | – | – |
MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers
|
sense-x/mixmim |
2022 |
| 13 |
UniRepLKNet-S |
– | 53 | – | – | – | – | – | – |
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
|
ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension |
2023 |
| 14 |
MixMIM-B |
– | 52.2 | – | – | – | – | – | – |
MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers
|
sense-x/mixmim |
2022 |
| 15 |
UniRepLKNet-T |
– | 51.7 | – | – | – | – | – | – |
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
|
ailab-cvc/unireplknet · Westlake-AI/openmixup · chenller/mmseg-extension |
2023 |
| 16 |
BiFormer-B (IN1k pretrain, MaskRCNN 12ep) |
– | 48.6 | – | – | – | – | – | – |
BiFormer: Vision Transformer with Bi-Level Routing Attention
|
rayleizhu/biformer · chenller/mmseg-extension · birder/birder |
2023 |
| 17 |
DeBiFormer-B (IN1k pretrain, MaskRCNN 12ep) |
– | 48.5 | – | – | – | – | – | – |
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
|
maclong01/DeBiFormer |
2024 |
| 18 |
BiFormer-S (IN1k pretrain, MaskRCNN 12ep) |
– | 47.8 | – | – | – | – | – | – |
BiFormer: Vision Transformer with Bi-Level Routing Attention
|
rayleizhu/biformer · chenller/mmseg-extension · birder/birder |
2023 |
| 19 |
DeBiFormer-S (IN1k pretrain, MaskRCNN 12ep) |
– | 47.5 | – | – | – | – | – | – |
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
|
maclong01/DeBiFormer |
2024 |
| 20 |
DeBiFormer-B (IN1k pretrain, Retina) |
– | 47.1 | – | – | – | – | – | – |
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
|
maclong01/DeBiFormer |
2024 |