paper-with-me

Video-based Generative Performance Benchmarking (Temporal Understanding) 벤치마크

Video-based Generative Performance Benchmarking (Temporal Understanding) on VideoInstruct

54개 결과 · ⬇ CSV · JSON

gpt-score

1.82 2.167 2.515 2.862 3.21 2023-04 2026-09 LLaMA Adapter — 1.98 (2023-04-28) LLaMA Adapter — 1.98 (2023-04-28) LLaMA Adapter — 1.98 (2023-04-28) Video Chat — 1.94 (2023-05-10) Video Chat — 1.94 (2023-05-10) Video Chat — 1.94 (2023-05-10) Video LLaMA — 1.82 (2023-06-05) Video LLaMA — 1.82 (2023-06-05) Video LLaMA — 1.82 (2023-06-05) Video-ChatGPT — 1.98 (2023-06-08) Video-ChatGPT — 1.98 (2023-06-08) Video-ChatGPT — 1.98 (2023-06-08) MovieChat — 2.24 (2023-07-31) MovieChat — 2.24 (2023-07-31) MovieChat — 2.24 (2023-07-31) BT-Adapter — 2.34 (2023-09-27) BT-Adapter (zero-shot) — 2.13 (2023-09-27) BT-Adapter — 2.34 (2023-09-27) BT-Adapter (zero-shot) — 2.13 (2023-09-27) BT-Adapter — 2.34 (2023-09-27) BT-Adapter (zero-shot) — 2.13 (2023-09-27) Chat-UniVi — 2.39 (2023-11-14) Chat-UniVi — 2.39 (2023-11-14) Chat-UniVi — 2.39 (2023-11-14) VideoChat2 — 2.66 (2023-11-28) VideoChat2_HD_mistral — 2.65 (2023-11-28) VideoChat2 — 2.66 (2023-11-28) VideoChat2_HD_mistral — 2.65 (2023-11-28) VideoChat2 — 2.66 (2023-11-28) VideoChat2_HD_mistral — 2.65 (2023-11-28) VTimeLLM — 2.49 (2023-11-30) VTimeLLM — 2.49 (2023-11-30) VTimeLLM — 2.49 (2023-11-30) ST-LLM — 2.93 (2024-03-30) ST-LLM — 2.93 (2024-03-30) ST-LLM — 2.93 (2024-03-30) MiniGPT4-video-7B — 2.65 (2024-04-04) MiniGPT4-video-7B — 2.65 (2024-04-04) MiniGPT4-video-7B — 2.65 (2024-04-04) PLLaVA-34B — 2.67 (2024-04-25) PLLaVA-34B — 2.67 (2024-04-25) PLLaVA-34B — 2.67 (2024-04-25) VideoGPT+ — 2.83 (2024-06-13) VideoGPT+ — 2.83 (2024-06-13) VideoGPT+ — 2.83 (2024-06-13) SlowFast-LLaVA-34B — 2.77 (2024-07-22) SlowFast-LLaVA-34B — 2.77 (2024-07-22) SlowFast-LLaVA-34B — 2.77 (2024-07-22) PPLLaVA-7B — 3.21 (2024-11-04) PPLLaVA-7B — 3.21 (2024-11-04) PPLLaVA-7B — 3.21 (2024-11-04) TS-LLaVA-34B — 2.77 (2024-11-17) TS-LLaVA-34B — 2.77 (2024-11-17) TS-LLaVA-34B — 2.77 (2024-11-17) LLaMA Adapter — 1.98 (2023-04-28) MovieChat — 2.24 (2023-07-31) BT-Adapter — 2.34 (2023-09-27) Chat-UniVi — 2.39 (2023-11-14) VideoChat2 — 2.66 (2023-11-28) ST-LLM — 2.93 (2024-03-30) PPLLaVA-7B — 3.21 (2024-11-04)
RankModel gpt-score PaperCodeYear
1 PPLLaVA-7B 3.21 PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance farewellthree/ppllava 2024
2 ST-LLM 2.93 ST-LLM: Large Language Models Are Effective Temporal Learners TencentARC/ST-LLM 2024
3 VideoGPT+ 2.83 VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding mbzuai-oryx/videogpt-plus 2024
4 SlowFast-LLaVA-34B 2.77 SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models apple/ml-slowfast-llava 2024
4 TS-LLaVA-34B 2.77 TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models tingyu215/ts-llava 2024
6 PLLaVA-34B 2.67 PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning magic-research/PLLaVA 2024
7 VideoChat2 2.66 MVBench: A Comprehensive Multi-modal Video Understanding Benchmark opengvlab/ask-anything · magic-research/PLLaVA · bytedance/tarsier 2023
8 MiniGPT4-video-7B 2.65 MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens Vision-CAIR/MiniGPT4-video · pwc-1/Paper-9 2024
8 VideoChat2_HD_mistral 2.65 MVBench: A Comprehensive Multi-modal Video Understanding Benchmark opengvlab/ask-anything · magic-research/PLLaVA · bytedance/tarsier 2023
10 VTimeLLM 2.49 VTimeLLM: Empower LLM to Grasp Video Moments huangb23/vtimellm 2023
11 Chat-UniVi 2.39 Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding pku-yuangroup/chat-univi · skyworkai/moh · skyworkai/moe-plus-plus · +1 2023
12 BT-Adapter 2.34 BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning farewellthree/BT-Adapter 2023
13 MovieChat 2.24 MovieChat: From Dense Token to Sparse Memory for Long Video Understanding rese1f/MovieChat 2023
14 BT-Adapter (zero-shot) 2.13 BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning farewellthree/BT-Adapter 2023
15 Video-ChatGPT 1.98 Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models mbzuai-oryx/video-chatgpt · qiujihao19/artemis 2023
15 LLaMA Adapter 1.98 LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model opengvlab/llama-adapter · zrrskywalker/llama-adapter · Mind23-2/MindCode-140 2023
17 Video Chat 1.94 VideoChat: Chat-Centric Video Understanding opengvlab/ask-anything 2023
18 Video LLaMA 1.82 Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding damo-nlp-sg/video-llama · damo-nlp-sg/videollama2 · damo-nlp-sg/videollama3 · +1 2023
19 PPLLaVA-7B 3.21 PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance farewellthree/ppllava 2024
20 ST-LLM 2.93 ST-LLM: Large Language Models Are Effective Temporal Learners TencentARC/ST-LLM 2024
1–20 / 54 다음 → 페이지당 10 20 50 100