paper-with-me

Zero-Shot Video Question Answer 벤치마크

Zero-Shot Video Question Answer on STAR Benchmark

8개 결과 · ⬇ CSV · JSON

Accuracy

41.6 45.95 50.3 54.65 59 2022-04 2026-09 Flamingo-9B — 41.8 (2022-04-29) Flamingo-9B — 41.8 (2022-04-29) InternVideo — 41.6 (2022-12-06) InternVideo — 41.6 (2022-12-06) VideoChat2 — 59.0 (2023-11-28) VideoChat2 — 59.0 (2023-11-28) VidCtx (7B) — 51.1 (2024-12-23) VidCtx (7B) — 51.1 (2024-12-23) Flamingo-9B — 41.8 (2022-04-29) VideoChat2 — 59.0 (2023-11-28)
RankModel Accuracy PaperCodeYear
1 VideoChat2 59.0 MVBench: A Comprehensive Multi-modal Video Understanding Benchmark opengvlab/ask-anything · magic-research/PLLaVA · bytedance/tarsier 2023
2 VidCtx (7B) 51.1 VidCtx: Context-aware Video Question Answering with Image Models idt-iti/vidctx 2024
3 Flamingo-9B 41.8 Flamingo: a Visual Language Model for Few-Shot Learning mlfoundations/open_flamingo · lucidrains/flamingo-pytorch · unispac/visual-adversarial-examples-jailbreak-large-language-models · +2 2022
4 InternVideo 41.6 InternVideo: General Video Foundation Models via Generative and Discriminative Learning opengvlab/internvideo · yingsen1/unimd 2022
5 VideoChat2 59.0 MVBench: A Comprehensive Multi-modal Video Understanding Benchmark opengvlab/ask-anything · magic-research/PLLaVA · bytedance/tarsier 2023
6 VidCtx (7B) 51.1 VidCtx: Context-aware Video Question Answering with Image Models idt-iti/vidctx 2024
7 Flamingo-9B 41.8 Flamingo: a Visual Language Model for Few-Shot Learning mlfoundations/open_flamingo · lucidrains/flamingo-pytorch · unispac/visual-adversarial-examples-jailbreak-large-language-models · +2 2022
8 InternVideo 41.6 InternVideo: General Video Foundation Models via Generative and Discriminative Learning opengvlab/internvideo · yingsen1/unimd 2022
1–8 / 8 페이지당 10 20 50 100