paper-with-me

Visual Reasoning 벤치마크

Visual Reasoning on NLVR2 Dev

15개 결과 · ⬇ CSV · JSON

Accuracy

66.7 72.9 79.11 85.31 91.51 2019-08 2026-09 VisualBERT — 66.7 (2019-08-09) LXMERT (Pre-train + scratch) — 74.9 (2019-08-20) ViLT-B/32 — 75.7 (2021-02-05) SOHO — 76.37 (2021-04-07) ALBEF (14M) — 83.14 (2021-07-16) SimVLM — 84.53 (2021-08-24) VLMo — 85.64 (2021-11-03) X-VLM (base) — 84.41 (2021-11-16) CoCa — 86.1 (2022-05-04) BEiT-3 — 91.51 (2022-08-22) X2-VLM (large) — 88.7 (2022-11-22) X2-VLM (base) — 86.2 (2022-11-22) XFM (base) — 87.6 (2023-01-12) VK-OOD — 83.9 (2023-02-11) VK-OOD — 84.6 (2023-09-21) VisualBERT — 66.7 (2019-08-09) LXMERT (Pre-train + scratch) — 74.9 (2019-08-20) ViLT-B/32 — 75.7 (2021-02-05) SOHO — 76.37 (2021-04-07) ALBEF (14M) — 83.14 (2021-07-16) SimVLM — 84.53 (2021-08-24) VLMo — 85.64 (2021-11-03) CoCa — 86.1 (2022-05-04) BEiT-3 — 91.51 (2022-08-22)
RankModel Accuracy PaperCodeYear
1 BEiT-3 91.51 Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks microsoft/unilm · lyan62/data-curation 2022
2 X2-VLM (large) 88.7 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
3 XFM (base) 87.6 Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks zhangxinsong-nlp/XFM 2023
4 X2-VLM (base) 86.2 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
5 CoCa 86.1 CoCa: Contrastive Captioners are Image-Text Foundation Models mlfoundations/open_clip · facebookresearch/multimodal · lucidrains/CoCa-pytorch · +3 2022
6 VLMo 85.64 VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts microsoft/unilm · ylsung/vl-merging 2021
7 VK-OOD 84.6 Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis ellenzhuwang/implicit_vkood 2023
8 SimVLM 84.53 SimVLM: Simple Visual Language Model Pretraining with Weak Supervision yulong-XJTU/SimVLM · FerryHuang/SimVLM 2021
9 X-VLM (base) 84.41 Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts zengyan-97/x-vlm 2021
10 VK-OOD 83.9 Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis ellenzhuwang/VK_OOD 2023
11 ALBEF (14M) 83.14 Align before Fuse: Vision and Language Representation Learning with Momentum Distillation salesforce/lavis · salesforce/ALBEF · facebookresearch/multimodal · +3 2021
12 SOHO 76.37 Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning researchmm/soho · PasserBy4/mypretraining · PasserBy4/pretraining 2021
13 ViLT-B/32 75.7 ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision huggingface/transformers · dandelin/vilt · glamor-usc/climb · +3 2021
14 LXMERT (Pre-train + scratch) 74.9 LXMERT: Learning Cross-Modality Encoder Representations from Transformers huggingface/transformers · airsplay/lxmert · zhegan27/VILLA · +6 2019
15 VisualBERT 66.7 VisualBERT: A Simple and Performant Baseline for Vision and Language uclanlp/visualbert · YIKUAN8/Transformers-VQA · lalithjets/surgical_vqa · +7 2019
1–15 / 15 페이지당 10 20 50 100