paper-with-me

Zero-Shot Cross-Modal Retrieval 벤치마크

Zero-Shot Cross-Modal Retrieval on Flickr30k

22개 결과 · ⬇ CSV · JSON

Image-to-text R@1

70.7 76.95 83.2 89.45 95.7 2019-09 2026-09 UNITER — 80.7 (2019-09-25) ImageBERT — 70.7 (2020-01-22) ViLT-B/32 — 73.2 (2021-02-05) ALIGN — 88.6 (2021-02-11) CLIP — 88.0 (2021-02-26) ALBEF — 90.5 (2021-07-16) Florence — 90.9 (2021-11-22) Flamingo — 89.3 (2022-04-29) CoCa — 92.5 (2022-05-04) BEiT-3 — 94.9 (2022-08-22) ERNIE-ViL 2.0 — 91.2 (2022-09-30) AltCLIP — 86.0 (2022-11-12) PTP-BLIP (14M) — 87.1 (2022-12-19) RO-ViT — 92.1 (2023-05-11) VK-OOD — 89.0 (2023-09-21) InternVL-G — 95.7 (2023-12-21) InternVL-C — 94.7 (2023-12-21) M2-Encoder — 91.2 (2024-01-29) COSMOS ViT-B/16 — 92.9 (2024-12-02) COSMOS ViT-B/32 — 89.9 (2024-12-02) UNITER — 80.7 (2019-09-25) ALIGN — 88.6 (2021-02-11) ALBEF — 90.5 (2021-07-16) Florence — 90.9 (2021-11-22) CoCa — 92.5 (2022-05-04) BEiT-3 — 94.9 (2022-08-22) InternVL-G — 95.7 (2023-12-21)
RankModel Image-to-text R@1Image-to-text R@5Image-to-text R@10Text-to-image R@1Text-to-image R@5Text-to-image R@10 Extra Training Data PaperCodeYear
1 InternVL-G 95.799.799.985.097.098.6 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks opengvlab/internvl · opengvlab/internvl-mmdetseg 2023
2 BEiT-3 94.999.9100.081.595.697.8 Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks microsoft/unilm · lyan62/data-curation 2022
3 InternVL-C 94.799.699.981.796.098.2 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks opengvlab/internvl · opengvlab/internvl-mmdetseg 2023
4 COSMOS ViT-B/16 92.999.499.980.395.397.6 COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training ExplainableML/cosmos 2024
5 CoCa 92.599.599.980.495.797.7 CoCa: Contrastive Captioners are Image-Text Foundation Models mlfoundations/open_clip · facebookresearch/multimodal · lucidrains/CoCa-pytorch · +3 2022
6 RO-ViT 92.199.499.780.796.197.7 Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers google-research/google-research · mcahny/rovit 2023
7 M2-Encoder 91.299.299.692.299.599.7 M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining alipay/Ant-Multi-Modal-Framework 2024
7 ERNIE-ViL 2.0 91.299.199.877.493.896.4 ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training PaddlePaddle/ERNIE 2022
9 Florence 90.999.1-76.793.6- Florence: A New Foundation Model for Computer Vision microsoft/unicl · MindCode-4/code-3 2021
10 ALBEF 90.598.899.776.893.796.7 Align before Fuse: Vision and Language Representation Learning with Momentum Distillation salesforce/lavis · salesforce/ALBEF · facebookresearch/multimodal · +3 2021
11 COSMOS ViT-B/32 89.998.899.376.192.896.2 COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training ExplainableML/cosmos 2024
12 Flamingo 89.398.899.779.595.397.9 Flamingo: a Visual Language Model for Few-Shot Learning mlfoundations/open_flamingo · lucidrains/flamingo-pytorch · unispac/visual-adversarial-examples-jailbreak-large-language-models · +2 2022
13 VK-OOD 89.099.299.877.294.398.2 Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis ellenzhuwang/implicit_vkood 2023
14 ALIGN 88.698.799.775.793.896.8 Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision facebookresearch/metaclip · kakaobrain/coyo-dataset · MicPie/clasp · +2 2021
15 CLIP 88.098.799.468.790.695.2 Learning Transferable Visual Models From Natural Language Supervision openai/CLIP · mlfoundations/open_clip · towhee-io/towhee · +79 2021
16 PTP-BLIP (14M) 87.198.499.373.191.094.8 Position-guided Text Prompt for Vision-Language Pre-training sail-sg/ptp 2022
17 AltCLIP 869899.172.591.695.4 AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities flagai-open/flagai · pwc-1/Paper-8 2022
18 UNITER 80.795.798.066.288.492.9 UNITER: UNiversal Image-TExt Representation Learning ChenRocks/UNITER · YIKUAN8/Transformers-VQA · necla-ml/SNLI-VE · +4 2019
19 ViLT-B/32 73.293.696.55582.589.8 ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision huggingface/transformers · dandelin/vilt · glamor-usc/climb · +3 2021
20 ImageBERT 70.790.294.054.379.687.5 ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data 2020
1–20 / 22 다음 → 페이지당 10 20 50 100