paper
-with-
me
Papers
Browse State-of-the-Art
Datasets
Methods
AI Agents
Trends
Digest
🌙
Visual Question Answering
벤치마크
Visual Question Answering on
BenchLMM
20개 결과 ·
⬇ CSV
·
JSON
GPT-3.5 score
30.1
37.17
44.23
51.3
58.37
2023-03
2026-09
GPT-4V — 58.37 (2023-03-15)
GPT-4V — 58.37 (2023-03-15)
LLaVA-1.5-7B — 46.83 (2023-04-17)
LLaVA-1-13B — 43.5 (2023-04-17)
LLaVA-1.5-7B — 46.83 (2023-04-17)
LLaVA-1-13B — 43.5 (2023-04-17)
MiniGPT4-13B — 34.93 (2023-04-20)
MiniGPT4-13B — 34.93 (2023-04-20)
Otter-7B — 39.13 (2023-05-05)
Otter-7B — 39.13 (2023-05-05)
InstructBLIP-13B — 45.03 (2023-05-11)
InstructBLIP-7B — 44.63 (2023-05-11)
InstructBLIP-13B — 45.03 (2023-05-11)
InstructBLIP-7B — 44.63 (2023-05-11)
LLaVA-1.5-13B — 55.53 (2023-10-05)
LLaVA-1.5-13B — 55.53 (2023-10-05)
MiniGPTv2-7B — 30.1 (2023-10-14)
MiniGPTv2-7B — 30.1 (2023-10-14)
Sphinx-V2-1K — 57.43 (2023-11-13)
Sphinx-V2-1K — 57.43 (2023-11-13)
GPT-4V — 58.37 (2023-03-15)
2023-03-15 — GPT-4V: GPT-3.5 score 58.37
Rank
Model
GPT-3.5 score
Extra Training Data
Paper
Code
Year
1
GPT-4V
58.37
✓
GPT-4 Technical Report
openai/evals
·
shmsw25/factscore
·
unispac/visual-adversarial-examples-jailbreak-large-language-models
·
+8
2023
2
Sphinx-V2-1K
57.43
✓
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
alpha-vllm/llama2-accessory
2023
3
LLaVA-1.5-13B
55.53
Improved Baselines with Visual Instruction Tuning
huggingface/transformers
·
haotian-liu/LLaVA
·
LLaVA-VL/LLaVA-NeXT
·
+6
2023
4
LLaVA-1.5-7B
46.83
Visual Instruction Tuning
huggingface/transformers
·
haotian-liu/LLaVA
·
LLaVA-VL/LLaVA-NeXT
·
+10
2023
5
InstructBLIP-13B
45.03
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
salesforce/lavis
·
tabtoyou/kollava
·
pwc-1/Paper-9
·
+1
2023
6
InstructBLIP-7B
44.63
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
salesforce/lavis
·
tabtoyou/kollava
·
pwc-1/Paper-9
·
+1
2023
7
LLaVA-1-13B
43.50
Visual Instruction Tuning
huggingface/transformers
·
haotian-liu/LLaVA
·
LLaVA-VL/LLaVA-NeXT
·
+10
2023
8
Otter-7B
39.13
Otter: A Multi-Modal Model with In-Context Instruction Tuning
luodian/otter
2023
9
MiniGPT4-13B
34.93
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
vision-cair/minigpt-4
·
zyang1580/binllm
·
2024-MindSpore-1/Code6
·
+3
2023
10
MiniGPTv2-7B
30.1
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
vision-cair/minigpt-4
·
zebangcheng/emotion-llama
2023
11
GPT-4V
58.37
✓
GPT-4 Technical Report
openai/evals
·
shmsw25/factscore
·
unispac/visual-adversarial-examples-jailbreak-large-language-models
·
+8
2023
12
Sphinx-V2-1K
57.43
✓
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
alpha-vllm/llama2-accessory
2023
13
LLaVA-1.5-13B
55.53
Improved Baselines with Visual Instruction Tuning
huggingface/transformers
·
haotian-liu/LLaVA
·
LLaVA-VL/LLaVA-NeXT
·
+6
2023
14
LLaVA-1.5-7B
46.83
Visual Instruction Tuning
huggingface/transformers
·
haotian-liu/LLaVA
·
LLaVA-VL/LLaVA-NeXT
·
+10
2023
15
InstructBLIP-13B
45.03
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
salesforce/lavis
·
tabtoyou/kollava
·
pwc-1/Paper-9
·
+1
2023
16
InstructBLIP-7B
44.63
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
salesforce/lavis
·
tabtoyou/kollava
·
pwc-1/Paper-9
·
+1
2023
17
LLaVA-1-13B
43.50
Visual Instruction Tuning
huggingface/transformers
·
haotian-liu/LLaVA
·
LLaVA-VL/LLaVA-NeXT
·
+10
2023
18
Otter-7B
39.13
Otter: A Multi-Modal Model with In-Context Instruction Tuning
luodian/otter
2023
19
MiniGPT4-13B
34.93
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
vision-cair/minigpt-4
·
zyang1580/binllm
·
2024-MindSpore-1/Code6
·
+3
2023
20
MiniGPTv2-7B
30.1
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
vision-cair/minigpt-4
·
zebangcheng/emotion-llama
2023
1–20 / 20
페이지당
10
20
50
100