| Rank | Model |
1:1 Accuracy |
Extra Training Data |
Paper | Code | Year |
| 1 |
ChartPaLI-5B + PaLM 2-S |
81.3 |
✓ |
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
|
|
2024 |
| 2 |
Gemini Ultra |
80.8 |
|
Gemini: A Family of Highly Capable Multimodal Models
|
valdecy/pybibx |
2023 |
| 3 |
DePlot+FlanPaLM+Codex (PoT Self-Consistency) |
79.3 |
|
DePlot: One-shot visual language reasoning by plot-to-table translation
|
huggingface/transformers |
2022 |
| 4 |
ChartPaLI-5B |
77.3 |
✓ |
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
|
|
2024 |
| 5 |
DePlot+Codex (PoT Self-Consistency) |
76.7 |
|
DePlot: One-shot visual language reasoning by plot-to-table translation
|
huggingface/transformers |
2022 |
| 5 |
ScreenAI 5B (4.62 B params, w/ OCR) |
76.7 |
✓ |
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
|
google-research-datasets/screen_qa · google-research-datasets/screen_annotation |
2024 |
| 7 |
SMoLA-PaLI-X Specialist Model |
74.6 |
✓ |
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts
|
|
2023 |
| 8 |
SMoLA-PaLI-X Generalist Model |
73.8 |
✓ |
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts
|
|
2023 |
| 9 |
MatCha4096 + LaMenDa |
72.64 |
✓ |
Synthesize Step-by-Step: Tools Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
|
|
2024 |
| 10 |
PaLI-X (Single-task FT w/ OCR) |
72.3 |
✓ |
PaLI-X: On Scaling up a Multilingual Vision and Language Model
|
kyegomez/PALI · doc-doc/NExT-OE |
2023 |
| 11 |
PaLI-X (Single-task FT) |
70.9 |
✓ |
PaLI-X: On Scaling up a Multilingual Vision and Language Model
|
kyegomez/PALI · doc-doc/NExT-OE |
2023 |
| 12 |
PaLI-X (Multi-task FT) |
70.6 |
✓ |
PaLI-X: On Scaling up a Multilingual Vision and Language Model
|
kyegomez/PALI · doc-doc/NExT-OE |
2023 |
| 13 |
DePlot+FlanPaLM (Self-Consistency) |
70.5 |
|
DePlot: One-shot visual language reasoning by plot-to-table translation
|
huggingface/transformers |
2022 |
| 14 |
PaLI-3 |
70 |
|
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
|
kyegomez/PALI3 |
2023 |
| 15 |
PaLI-3 (w/ OCR) |
69.5 |
|
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
|
kyegomez/PALI3 |
2023 |
| 16 |
DePlot+FlanPaLM (CoT) |
67.3 |
|
DePlot: One-shot visual language reasoning by plot-to-table translation
|
huggingface/transformers |
2022 |
| 17 |
Qwen-VL-Chat |
66.3 |
✓ |
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
|
qwenlm/qwen-vl · brandon3964/multimodal-task-vector |
2023 |
| 18 |
UniChart |
66.24 |
✓ |
UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
|
vis-nlp/unichart |
2023 |
| 19 |
Qwen-VL |
65.7 |
✓ |
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
|
qwenlm/qwen-vl · brandon3964/multimodal-task-vector |
2023 |
| 20 |
StructChart+GPT3.5 (STR ChartQA+SimChart9K) |
65.3 |
✓ |
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
|
alpha-innovator/chartvlm · unimodal4reasoning/chartvlm · unimodal4reasoning/simchart9k |
2023 |