paper-with-me

Papers

Evaluating Graphical Perception Capabilities of Vision Transformers

2026-02-20 · Poonam Poonam, Pere-Pau Vázquez, Timo Ropinski arxiv

Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphical perception tasks, which are essential for interpreting visualizations, the perceptual capabilities of ViTs remain largely unexplored. In this work, we investigate the performance of ViTs in elementary visual judgment tasks inspired by the foundational studies of Cleveland and McGill, which quantified the accuracy of human perception across different visual encodings. Inspired by their study, we benchmark ViTs against CNNs and human participants in a series of controlled graphical perception tasks. Our results reveal that, although ViTs demonstrate strong performance in general vision tasks, their alignment with human-like graphical perception in the visualization domain is limited. This study highlights key perceptual gaps and points to important considerations for the application of ViTs in visualization systems and graphical perceptual modeling.

📄 PDF Abstract BibTeX arXiv:2602.18178

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models

2024-10-31 · Grace Guo, Jenna Jiayi Kang, Raj Sanjay Shah, Hanspeter Pfister 외

Vision Language Models (VLMs) have been successful at many chart comprehension tasks that require attending to both the images of charts and their accompanying textual descriptions. However, it is not well established ho…

Data Visualization

Towards Understanding Graphical Perception in Large Multimodal Models

2025-03-13 · Kai Zhang, Jianwei Yang, Jeevana Priya Inala, Chandan Singh 외

Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprisingly found that these models struggle with…

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

2026-06-12 · Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski 외 arxiv

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular appli…

Do Vision Transformers See Like Humans? Evaluating their Perceptual Alignment

2025-08-13 · Pablo Hernández-Cámara, Jose Manuel Jaén-Lorites, Jorge Vila-Tomás, Valero Laparra 외 arxiv

Vision Transformers (ViTs) achieve remarkable performance in image recognition tasks, yet their alignment with human perception remains largely unexplored. This study systematically analyzes how model size, dataset size,…

Data Augmentation

Assessing Graphical Perception of Image Embedding Models using Channel Effectiveness

2024-07-30 · Soohyun Lee, Minsuk Chang, SeokHyeon Park, Jinwook Seo

Recent advancements in vision models have greatly improved their ability to handle complex chart understanding tasks, like chart captioning and question answering. However, it remains challenging to assess how these mode…

Chart UnderstandingQuestion Answering