paper-with-me

Zero-Shot Video Question Answer 벤치마크

Zero-Shot Video Question Answer on Video-MME (w/o subs)

18개 결과 · ⬇ CSV · JSON

Accuracy (%)

46.3 54.08 61.85 69.62 77.4 2023-12 2026-09 VILA-1.5 (34B) — 61.4 (2023-12-12) VILA-1.5 (34B) — 61.4 (2023-12-12) Gemini 1.5 Pro — 71.9 (2024-03-08) Gemini 1.5 Flash — 66.3 (2024-03-08) Gemini 1.5 Pro — 71.9 (2024-03-08) Gemini 1.5 Flash — 66.3 (2024-03-08) VideoLLaMA2 (72B) — 60.9 (2024-06-11) VideoLLaMA2 (72B) — 60.9 (2024-06-11) GPT-4o — 70.3 (2024-06-14) GPT-4o mini — 62.3 (2024-06-14) GPT-4o — 70.3 (2024-06-14) GPT-4o mini — 62.3 (2024-06-14) VideoChat-T (7B) — 46.3 (2024-10-25) VideoChat-T (7B) — 46.3 (2024-10-25) Video-RAG (based on LLaVA-Video) — 77.4 (2024-11-20) Video-RAG (based on LLaVA-Video) — 77.4 (2024-11-20) VILA-1.5 (34B) — 61.4 (2023-12-12) Gemini 1.5 Pro — 71.9 (2024-03-08) Video-RAG (based on LLaVA-Video) — 77.4 (2024-11-20)
RankModel Accuracy (%) PaperCodeYear
1 Video-RAG (based on LLaVA-Video) 77.4 Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension leon1207/video-rag-master 2024
2 Gemini 1.5 Pro 71.9 Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context dlvuldet/primevul 2024
3 GPT-4o 70.3 GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding 2024
4 Gemini 1.5 Flash 66.3 Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context dlvuldet/primevul 2024
5 LLaVA-OneVision (72B) 64.8
6 GPT-4o mini 62.3 GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding 2024
7 VILA-1.5 (34B) 61.4 VILA: On Pre-training for Visual Language Models efficient-large-model/vila · nvlabs/vila · mit-han-lab/llm-awq 2023
8 VideoLLaMA2 (72B) 60.9 VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs damo-nlp-sg/videollama2 · damo-nlp-sg/videollama3 · damo-nlp-sg/inf-clip 2024
9 VideoChat-T (7B) 46.3 TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning OpenGVLab/TimeSuite 2024
10 Video-RAG (based on LLaVA-Video) 77.4 Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension leon1207/video-rag-master 2024
11 Gemini 1.5 Pro 71.9 Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context dlvuldet/primevul 2024
12 GPT-4o 70.3 GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding 2024
13 Gemini 1.5 Flash 66.3 Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context dlvuldet/primevul 2024
14 LLaVA-OneVision (72B) 64.8
15 GPT-4o mini 62.3 GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding 2024
16 VILA-1.5 (34B) 61.4 VILA: On Pre-training for Visual Language Models efficient-large-model/vila · nvlabs/vila · mit-han-lab/llm-awq 2023
17 VideoLLaMA2 (72B) 60.9 VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs damo-nlp-sg/videollama2 · damo-nlp-sg/videollama3 · damo-nlp-sg/inf-clip 2024
18 VideoChat-T (7B) 46.3 TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning OpenGVLab/TimeSuite 2024
1–18 / 18 페이지당 10 20 50 100