paper-with-me

Visual Question Answering 벤치마크

Visual Question Answering on BenchLMM

20개 결과 · ⬇ CSV · JSON

GPT-3.5 score

30.1 37.17 44.23 51.3 58.37 2023-03 2026-09 GPT-4V — 58.37 (2023-03-15) GPT-4V — 58.37 (2023-03-15) LLaVA-1.5-7B — 46.83 (2023-04-17) LLaVA-1-13B — 43.5 (2023-04-17) LLaVA-1.5-7B — 46.83 (2023-04-17) LLaVA-1-13B — 43.5 (2023-04-17) MiniGPT4-13B — 34.93 (2023-04-20) MiniGPT4-13B — 34.93 (2023-04-20) Otter-7B — 39.13 (2023-05-05) Otter-7B — 39.13 (2023-05-05) InstructBLIP-13B — 45.03 (2023-05-11) InstructBLIP-7B — 44.63 (2023-05-11) InstructBLIP-13B — 45.03 (2023-05-11) InstructBLIP-7B — 44.63 (2023-05-11) LLaVA-1.5-13B — 55.53 (2023-10-05) LLaVA-1.5-13B — 55.53 (2023-10-05) MiniGPTv2-7B — 30.1 (2023-10-14) MiniGPTv2-7B — 30.1 (2023-10-14) Sphinx-V2-1K — 57.43 (2023-11-13) Sphinx-V2-1K — 57.43 (2023-11-13) GPT-4V — 58.37 (2023-03-15)
RankModel GPT-3.5 score Extra Training Data PaperCodeYear
1 GPT-4V 58.37 GPT-4 Technical Report openai/evals · shmsw25/factscore · unispac/visual-adversarial-examples-jailbreak-large-language-models · +8 2023
2 Sphinx-V2-1K 57.43 SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models alpha-vllm/llama2-accessory 2023
3 LLaVA-1.5-13B 55.53 Improved Baselines with Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +6 2023
4 LLaVA-1.5-7B 46.83 Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +10 2023
5 InstructBLIP-13B 45.03 InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning salesforce/lavis · tabtoyou/kollava · pwc-1/Paper-9 · +1 2023
6 InstructBLIP-7B 44.63 InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning salesforce/lavis · tabtoyou/kollava · pwc-1/Paper-9 · +1 2023
7 LLaVA-1-13B 43.50 Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +10 2023
8 Otter-7B 39.13 Otter: A Multi-Modal Model with In-Context Instruction Tuning luodian/otter 2023
9 MiniGPT4-13B 34.93 MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models vision-cair/minigpt-4 · zyang1580/binllm · 2024-MindSpore-1/Code6 · +3 2023
10 MiniGPTv2-7B 30.1 MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning vision-cair/minigpt-4 · zebangcheng/emotion-llama 2023
11 GPT-4V 58.37 GPT-4 Technical Report openai/evals · shmsw25/factscore · unispac/visual-adversarial-examples-jailbreak-large-language-models · +8 2023
12 Sphinx-V2-1K 57.43 SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models alpha-vllm/llama2-accessory 2023
13 LLaVA-1.5-13B 55.53 Improved Baselines with Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +6 2023
14 LLaVA-1.5-7B 46.83 Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +10 2023
15 InstructBLIP-13B 45.03 InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning salesforce/lavis · tabtoyou/kollava · pwc-1/Paper-9 · +1 2023
16 InstructBLIP-7B 44.63 InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning salesforce/lavis · tabtoyou/kollava · pwc-1/Paper-9 · +1 2023
17 LLaVA-1-13B 43.50 Visual Instruction Tuning huggingface/transformers · haotian-liu/LLaVA · LLaVA-VL/LLaVA-NeXT · +10 2023
18 Otter-7B 39.13 Otter: A Multi-Modal Model with In-Context Instruction Tuning luodian/otter 2023
19 MiniGPT4-13B 34.93 MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models vision-cair/minigpt-4 · zyang1580/binllm · 2024-MindSpore-1/Code6 · +3 2023
20 MiniGPTv2-7B 30.1 MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning vision-cair/minigpt-4 · zebangcheng/emotion-llama 2023
1–20 / 20 페이지당 10 20 50 100