Visual Question Answering (VQA) 벤치마크
Visual Question Answering (VQA) on AI2D
EM
- 2023-05-23 — DUBLIN: EM 51.11
- 2023-12-01 — SMoLA-PaLI-X Specialist Model: EM 82.5
| Rank | Model | EM | Extra Training Data | Paper | Code | Year |
|---|---|---|---|---|---|---|
| 1 | SMoLA-PaLI-X Specialist Model | 82.5 | ✓ | Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts | 2023 | |
| 2 | SMoLA-PaLI-X Generalist Model | 81.4 | ✓ | Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts | 2023 | |
| 3 | Gemini Ultra | 79.5 | Gemini: A Family of Highly Capable Multimodal Models | valdecy/pybibx | 2023 | |
| 4 | DUBLIN | 51.11 | DUBLIN -- Document Understanding By Language-Image Network | 2023 |