paper-with-me

홈 › Papers

EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts

2025-08-06 · Kushin Mukherjee, Donghao Ren, Dominik Moritz, Yannick Assogba arxiv

Multimodal vision-language models (VLMs) continue to achieve ever-improving scores on chart understanding benchmarks. Yet, we find that this progress does not fully capture the breadth of visual reasoning capabilities essential for interpreting charts. We introduce EncQA, a novel benchmark informed by the visualization literature, designed to provide systematic coverage of visual encodings and analytic tasks that are crucial for chart understanding. EncQA provides 2,076 synthetic question-answer pairs, enabling balanced coverage of six visual encoding channels (position, length, area, color quantitative, color nominal, and shape) and eight tasks (find extrema, retrieve value, find anomaly, filter values, compute derived value exact, compute derived value relative, correlate values, and correlate values relative). Our evaluation of 9 state-of-the-art VLMs reveals that performance varies significantly across encodings within the same task, as well as across tasks. Contrary to expectations, we observe that performance does not improve with model size for many task-encoding pairs. Our results suggest that advancing chart understanding requires targeted strategies addressing specific visual reasoning gaps, rather than solely scaling up model or dataset size.

📄 PDF Abstract BibTeX arXiv:2508.04650

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

OpenCQA: Open-ended Question Answering with Charts

2022-10-12 · Shankar Kantharaj, Xuan Long Do, Rixie Tiffany Ko Leong, Jia Qing Tan 외

Charts are very popular to analyze data and convey important insights. People often analyze visualizations to answer open-ended questions that require explanatory answers. Answering such questions are often difficult and…

Arithmetic ReasoningDescriptiveOpen-Ended Question AnsweringQuestion Answering

MSG-Chart: Multimodal Scene Graph for ChartQA

2024-08-09 · Yue Dai, Soyeon Caren Han, Wei Liu

Automatic Chart Question Answering (ChartQA) is challenging due to the complex distribution of chart elements with patterns of the underlying data not explicitly displayed in charts. To address this challenge, we design …

Chart Question AnsweringInductive BiasQuestion Answering

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

2026-06-09 · Sunil Khatri, Steven Landgraf, Markus Ulrich, Simon Reiß arxiv

Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at test-time. With the history of in-context learning in natural langua…

Lost in Translation: Do LVLM Judges Generalize Across Languages?

2026-04-21 · Md Tahmid Rahman Laskar, Mohammed Saidul Islam, Mir Tafseer Nayeem, Amran Bhuiyan 외 arxiv

Automatic evaluators such as reward models play a central role in the alignment and evaluation of large vision-language models (LVLMs). Despite their growing importance, these evaluators are almost exclusively assessed o…

Domain Adaptation

TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning

2024-04-25 · Liang Zhang, Anwen Hu, Haiyang Xu, Ming Yan 외

Charts are important for presenting and explaining complex data relationships. Recently, multimodal large language models (MLLMs) have shown remarkable capabilities in various chart understanding tasks. However, the shee…

Chart Understanding