paper-with-me

Visual Question Answering (VQA) 벤치마크

Visual Question Answering (VQA) on MSVD-QA

36개 결과 · ⬇ CSV · JSON

Accuracy

0.313 150.7 301.2 451.6 602 2017-04 2026-09 ST-VQA — 0.313 (2017-04-14) Co-Mem — 0.317 (2018-03-29) HMEMA — 0.337 (2019-04-08) HCRN — 0.361 (2020-02-25) SSML — 0.351 (2020-03-06) DualVGR — 0.39 (2021-07-10) ALPRO — 0.459 (2021-12-17) All-in-one-B — 0.483 (2022-03-14) GIT — 0.568 (2022-05-27) Clover — 0.524 (2022-07-16) Co-Tokenization — 486.0 (2022-08-01) VIOLETv2 — 0.547 (2022-09-04) OmniVL — 0.51 (2022-09-15) X2-VLM (large) — 0.546 (2022-11-22) X2-VLM (base) — 0.528 (2022-11-22) InternVideo — 0.555 (2022-12-06) VideoCoCa — 0.569 (2022-12-09) HiTeA — 0.556 (2022-12-30) mPLUG-2 — 0.581 (2023-02-01) MuLTI — 0.547 (2023-03-10) VIOLET + MELTR — 0.517 (2023-03-23) UMT-L (ViT-L/16) — 0.552 (2023-03-28) MaMMUT (ours) — 602.0 (2023-03-29) VALOR — 0.6 (2023-04-17) VLAB — 0.61 (2023-05-22) VAST — 0.6 (2023-05-29) COSA — 0.6 (2023-06-15) LRCE — 0.478 (2023-06-30) GIT+MDF — 0.469 (2023-07-09) AIO+MIF — 0.467 (2023-07-09) FrozenBiLM+ — 0.558 (2023-08-18) VIOLET+ — 0.495 (2023-08-18) JustAsk+ — 0.477 (2023-08-18) All-in-one+ — 0.438 (2023-08-18) vid-TLDR (UMT-L) — 0.549 (2024-03-20) MA-LMM — 0.606 (2024-04-08) ST-VQA — 0.313 (2017-04-14) Co-Mem — 0.317 (2018-03-29) HMEMA — 0.337 (2019-04-08) HCRN — 0.361 (2020-02-25) DualVGR — 0.39 (2021-07-10) ALPRO — 0.459 (2021-12-17) All-in-one-B — 0.483 (2022-03-14) GIT — 0.568 (2022-05-27) Co-Tokenization — 486.0 (2022-08-01) MaMMUT (ours) — 602.0 (2023-03-29)
RankModel Accuracy Extra Training Data PaperCodeYear
1 VLAB 0.61 VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending 2023
2 MA-LMM 0.606 MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding boheumd/MA-LMM 2024
3 MaMMUT (ours) .602 MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks lucidrains/mammut-pytorch 2023
4 VALOR 0.60 VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset TXH-mercury/VALOR 2023
4 VAST 0.60 VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset TXH-mercury/VALOR · txh-mercury/vast 2023
4 COSA 0.60 COSA: Concatenated Sample Pretrained Vision-Language Foundation Model txh-mercury/cosa 2023
7 mPLUG-2 0.581 mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video modelscope/modelscope · x-plug/mplug-owl · alibaba/AliceMind · +1 2023
8 VideoCoCa 0.569 VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners 2022
9 GIT 0.568 GIT: A Generative Image-to-text Transformer for Vision and Language microsoft/GenerativeImage2Text 2022
10 FrozenBiLM+ 0.558 Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models mlvlab/ovqa 2023
1–10 / 36 다음 → 페이지당 10 20 50 100