paper-with-me

Visual Question Answering (VQA) 벤치마크

Visual Question Answering (VQA) on MSVD-QA

36개 결과 · ⬇ CSV · JSON

Accuracy

0.313 150.7 301.2 451.6 602 2017-04 2026-09 ST-VQA — 0.313 (2017-04-14) Co-Mem — 0.317 (2018-03-29) HMEMA — 0.337 (2019-04-08) HCRN — 0.361 (2020-02-25) SSML — 0.351 (2020-03-06) DualVGR — 0.39 (2021-07-10) ALPRO — 0.459 (2021-12-17) All-in-one-B — 0.483 (2022-03-14) GIT — 0.568 (2022-05-27) Clover — 0.524 (2022-07-16) Co-Tokenization — 486.0 (2022-08-01) VIOLETv2 — 0.547 (2022-09-04) OmniVL — 0.51 (2022-09-15) X2-VLM (large) — 0.546 (2022-11-22) X2-VLM (base) — 0.528 (2022-11-22) InternVideo — 0.555 (2022-12-06) VideoCoCa — 0.569 (2022-12-09) HiTeA — 0.556 (2022-12-30) mPLUG-2 — 0.581 (2023-02-01) MuLTI — 0.547 (2023-03-10) VIOLET + MELTR — 0.517 (2023-03-23) UMT-L (ViT-L/16) — 0.552 (2023-03-28) MaMMUT (ours) — 602.0 (2023-03-29) VALOR — 0.6 (2023-04-17) VLAB — 0.61 (2023-05-22) VAST — 0.6 (2023-05-29) COSA — 0.6 (2023-06-15) LRCE — 0.478 (2023-06-30) GIT+MDF — 0.469 (2023-07-09) AIO+MIF — 0.467 (2023-07-09) FrozenBiLM+ — 0.558 (2023-08-18) VIOLET+ — 0.495 (2023-08-18) JustAsk+ — 0.477 (2023-08-18) All-in-one+ — 0.438 (2023-08-18) vid-TLDR (UMT-L) — 0.549 (2024-03-20) MA-LMM — 0.606 (2024-04-08) ST-VQA — 0.313 (2017-04-14) Co-Mem — 0.317 (2018-03-29) HMEMA — 0.337 (2019-04-08) HCRN — 0.361 (2020-02-25) DualVGR — 0.39 (2021-07-10) ALPRO — 0.459 (2021-12-17) All-in-one-B — 0.483 (2022-03-14) GIT — 0.568 (2022-05-27) Co-Tokenization — 486.0 (2022-08-01) MaMMUT (ours) — 602.0 (2023-03-29)
RankModel Accuracy Extra Training Data PaperCodeYear
1 VLAB 0.61 VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending 2023
2 MA-LMM 0.606 MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding boheumd/MA-LMM 2024
3 MaMMUT (ours) .602 MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks lucidrains/mammut-pytorch 2023
4 VALOR 0.60 VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset TXH-mercury/VALOR 2023
4 VAST 0.60 VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset TXH-mercury/VALOR · txh-mercury/vast 2023
4 COSA 0.60 COSA: Concatenated Sample Pretrained Vision-Language Foundation Model txh-mercury/cosa 2023
7 mPLUG-2 0.581 mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video modelscope/modelscope · x-plug/mplug-owl · alibaba/AliceMind · +1 2023
8 VideoCoCa 0.569 VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners 2022
9 GIT 0.568 GIT: A Generative Image-to-text Transformer for Vision and Language microsoft/GenerativeImage2Text 2022
10 FrozenBiLM+ 0.558 Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models mlvlab/ovqa 2023
11 HiTeA 0.556 HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training 2022
12 InternVideo 0.555 InternVideo: General Video Foundation Models via Generative and Discriminative Learning opengvlab/internvideo · yingsen1/unimd 2022
13 UMT-L (ViT-L/16) 0.552 Unmasked Teacher: Towards Training-Efficient Video Foundation Models opengvlab/unmasked_teacher 2023
14 vid-TLDR (UMT-L) 0.549 vid-TLDR: Training Free Token merging for Light-weight Video Transformer mlvlab/vid-tldr 2024
15 VIOLETv2 0.547 An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling tsujuifu/pytorch_empirical-mvm 2022
15 MuLTI 0.547 MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling 2023
17 X2-VLM (large) 0.546 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
18 X2-VLM (base) 0.528 X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks zengyan-97/x-vlm · zengyan-97/x2-vlm 2022
19 Clover 0.524 Clover: Towards A Unified Video-Language Alignment and Fusion Model leeyn-43/clover 2022
20 VIOLET + MELTR 0.517 MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models mlvlab/MELTR 2023
21 OmniVL 0.510 OmniVL:One Foundation Model for Image-Language and Video-Language Tasks 2022
22 VIOLET+ 0.495 Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models mlvlab/ovqa 2023
23 Co-Tokenization .486 Video Question Answering with Iterative Video-Text Co-Tokenization 2022
24 All-in-one-B 0.483 All in One: Exploring Unified Video-Language Pre-training showlab/all-in-one 2022
25 LRCE 0.478 Lightweight Recurrent Cross-modal Encoder for Video Question Answering Sejong-VLI/VQA-LRCE-KBS-2023 2023
26 JustAsk+ 0.477 Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models mlvlab/ovqa 2023
27 GIT+MDF 0.469 Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models declare-lab/sealing · declare-lab/sas-vqa 2023
28 AIO+MIF 0.467 Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models declare-lab/sealing · declare-lab/sas-vqa 2023
29 ALPRO 0.459 Align and Prompt: Video-and-Language Pre-training with Entity Prompts salesforce/alpro 2021
30 All-in-one+ 0.438 Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models mlvlab/ovqa 2023
31 DualVGR 0.390 DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering MM-IR/DualVGR-VideoQA 2021
32 HCRN 0.361 Hierarchical Conditional Relation Networks for Video Question Answering thaolmk54/hcrn-videoqa 2020
33 SSML 0.351 Noise Estimation Using Density Estimation for Self-Supervised Multimodal Learning elad-amrani/ssml 2020
34 HMEMA 0.337 Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering fanchenyou/HME-VideoQA 2019
35 Co-Mem 0.317 Motion-Appearance Co-Memory Networks for Video Question Answering 2018
36 ST-VQA 0.313 TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering ahjeongseo/MASN-pytorch · chaitanyadwivedii/3D-Attention-is-All-You-Need 2017
1–36 / 36 페이지당 10 20 50 100