paper-with-me

Visual Grounding 벤치마크

Visual Grounding on RefCOCO+ testA

7개 결과 · ⬇ CSV · JSON

Accuracy (%)

89 90.58 92.15 93.72 95.3 2021-11 2026-09 X-VLM (base) — 89.0 (2021-11-16) X2-VLM (large) — 92.1 (2022-11-22) X2-VLM (base) — 90.3 (2022-11-22) XFM (base) — 90.4 (2023-01-12) mPLUG-2 — 92.8 (2023-02-01) Florence-2-large-ft — 95.3 (2023-11-10) X-VLM (base) — 89.0 (2021-11-16) X2-VLM (large) — 92.1 (2022-11-22) mPLUG-2 — 92.8 (2023-02-01) Florence-2-large-ft — 95.3 (2023-11-10)
RankModel Accuracy (%)IoU Extra Training Data PaperCodeYear
1 Florence-2-large-ft 95.3 Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks retkowsky/florence-2 2023
2 mPLUG-2 92.8 mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video modelscope/modelscope · x-plug/mplug-owl · alibaba/AliceMind · +1 2023
3 X2-VLM (large) 92.1 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
4 XFM (base) 90.4 Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks zhangxinsong-nlp/XFM 2023
5 X2-VLM (base) 90.3 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
6 X-VLM (base) 89.00 Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts zengyan-97/x-vlm 2021
7 HYDRA 61.1 HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning ControlNet/HYDRA 2024
1–7 / 7 페이지당 10 20 50 100