paper-with-me

Visual Reasoning 벤치마크

Visual Reasoning on NLVR2 Test

14개 결과 · ⬇ CSV · JSON

Accuracy

76.13 80.24 84.35 88.47 92.58 2019-08 2026-09 LXMERT — 76.2 (2019-08-20) UNITER (Large) — 79.5 (2019-09-25) ViLT-B/32 — 76.13 (2021-02-05) SOHO — 77.32 (2021-04-07) ALBEF (14M) — 82.55 (2021-07-16) SimVLM — 85.15 (2021-08-24) VLMo — 86.86 (2021-11-03) X-VLM (base) — 84.76 (2021-11-16) BLIP-129M — 83.09 (2022-01-28) CoCa — 87.0 (2022-05-04) BEiT-3 — 92.58 (2022-08-22) X2-VLM (large) — 89.4 (2022-11-22) X2-VLM (base) — 87.0 (2022-11-22) XFM (base) — 88.4 (2023-01-12) LXMERT — 76.2 (2019-08-20) UNITER (Large) — 79.5 (2019-09-25) ALBEF (14M) — 82.55 (2021-07-16) SimVLM — 85.15 (2021-08-24) VLMo — 86.86 (2021-11-03) CoCa — 87.0 (2022-05-04) BEiT-3 — 92.58 (2022-08-22)
RankModel Accuracy PaperCodeYear
1 BEiT-3 92.58 Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks microsoft/unilm · lyan62/data-curation 2022
2 X2-VLM (large) 89.4 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
3 XFM (base) 88.4 Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks zhangxinsong-nlp/XFM 2023
4 CoCa 87.0 CoCa: Contrastive Captioners are Image-Text Foundation Models mlfoundations/open_clip · facebookresearch/multimodal · lucidrains/CoCa-pytorch · +3 2022
4 X2-VLM (base) 87.0 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
6 VLMo 86.86 VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts microsoft/unilm · ylsung/vl-merging 2021
7 SimVLM 85.15 SimVLM: Simple Visual Language Model Pretraining with Weak Supervision yulong-XJTU/SimVLM · FerryHuang/SimVLM 2021
8 X-VLM (base) 84.76 Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts zengyan-97/x-vlm 2021
9 BLIP-129M 83.09 BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation huggingface/transformers · salesforce/lavis · salesforce/blip · +6 2022
10 ALBEF (14M) 82.55 Align before Fuse: Vision and Language Representation Learning with Momentum Distillation salesforce/lavis · salesforce/ALBEF · facebookresearch/multimodal · +3 2021
11 UNITER (Large) 79.5 UNITER: UNiversal Image-TExt Representation Learning ChenRocks/UNITER · YIKUAN8/Transformers-VQA · necla-ml/SNLI-VE · +4 2019
12 SOHO 77.32 Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning researchmm/soho · PasserBy4/mypretraining · PasserBy4/pretraining 2021
13 LXMERT 76.2 LXMERT: Learning Cross-Modality Encoder Representations from Transformers huggingface/transformers · airsplay/lxmert · zhegan27/VILLA · +6 2019
14 ViLT-B/32 76.13 ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision huggingface/transformers · dandelin/vilt · glamor-usc/climb · +3 2021
1–14 / 14 페이지당 10 20 50 100