| Rank | Model |
mAP |
Paper | Code | Year |
| 1 |
Diffusion-Based Cross-Modal Feature Extr
자동 추출 |
98.6 |
Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
|
|
2025 |
| 2 |
ADDS(ViT-L-336, resolution 1344) |
93.54 |
Open Vocabulary Multi-Label Classification with Dual-Modal Decoder on Aligned Visual-Textual Features
|
|
2022 |
| 3 |
ADDS(ViT-L-336, resolution 640) |
93.41 |
Open Vocabulary Multi-Label Classification with Dual-Modal Decoder on Aligned Visual-Textual Features
|
|
2022 |
| 4 |
ADDS(ViT-L-336, resolution 336) |
91.76 |
Open Vocabulary Multi-Label Classification with Dual-Modal Decoder on Aligned Visual-Textual Features
|
|
2022 |
| 5 |
ML-Decoder(TResNet-XL, resolution 640) |
91.4 |
ML-Decoder: Scalable and Versatile Classification Head
|
alibaba-miil/ml_decoder |
2021 |
| 6 |
Q2L-CvT(ImageNet-21K pretraining, resolution 384) |
91.3 |
Query2Label: A Simple Transformer Way to Multi-Label Classification
|
SlongLiu/query2labels · curt-tigges/query2label · averyfallson/rmffn |
2021 |
| 6 |
MLD-TResNet-L-AAM[640x640] |
91.30 |
Combining Metric Learning and Attention Heads For Accurate and Efficient Multilabel Image Classification
|
openvinotoolkit/deep-object-reid |
2022 |
| 8 |
ML-Decoder(TResNet-L, resolution 640) |
91.1 |
ML-Decoder: Scalable and Versatile Classification Head
|
alibaba-miil/ml_decoder |
2021 |
| 9 |
Q2L-SwinL(ImageNet-21K pretraining, resolution 384) |
90.5 |
Query2Label: A Simple Transformer Way to Multi-Label Classification
|
SlongLiu/query2labels · curt-tigges/query2label · averyfallson/rmffn |
2021 |
| 10 |
Q2L-TResL(ImageNet-21K pretraining, resolution 640) |
90.3 |
Query2Label: A Simple Transformer Way to Multi-Label Classification
|
SlongLiu/query2labels · curt-tigges/query2label · averyfallson/rmffn |
2021 |
| 10 |
IDA-SwinL |
90.3 |
Causality Compensated Attention for Contextual Biased Visual Recognition
|
yu-gi-oh-leilei/IDA_2023ICLR |
2023 |
| 10 |
CCD-SwinL |
90.3 |
Contextual Debiasing for Visual Recognition With Causal Mechanisms
|
farewellthree/Causal-Context-Debiasing |
2022 |
| 13 |
MlTr-XL(ImageNet-21K pretraining, resolution 384) |
90.0 |
MlTr: Multi-label Classification with Transformer
|
starmemda/MlTr |
2021 |
| 14 |
TResNet-L-V2, (ImageNet-21K-P pretraining, resolution 640) |
89.8 |
ImageNet-21K Pretraining for the Masses
|
Alibaba-MIIL/ImageNet21K · YutingLi0606/SURE · encounter1997/fp-detr
· +2 |
2021 |
| 15 |
MlTr-L(ImageNet-21K pretraining, resolution 384) |
88.5 |
MlTr: Multi-label Classification with Transformer
|
starmemda/MlTr |
2021 |
| 16 |
TResNet-XL (resolution 640) |
88.4 |
Asymmetric Loss For Multi-Label Classification
|
Alibaba-MIIL/ASL · mrT23/TResNet · Alibaba-MIIL/TResNet
· +2 |
2020 |
| 16 |
TResNet-L-V2, (ImageNet-21K-P pretraining, resolution 448) |
88.4 |
ImageNet-21K Pretraining for the Masses
|
Alibaba-MIIL/ImageNet21K · YutingLi0606/SURE · encounter1997/fp-detr
· +2 |
2021 |
| 18 |
GKGNet(resolution 576) |
87.7 |
GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
|
jin-s13/gkgnet |
2023 |
| 19 |
M3TR(ImageNet-21K-P pretraining, resolution 448) |
87.5 |
M3TR: Multi-modal Multi-label Recognition with Transformer
|
iCVTEAM/M3TR |
2021 |
| 20 |
GKGNet(resolution 448) |
86.7 |
GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
|
jin-s13/gkgnet |
2023 |