paper-with-me

홈 › Papers

TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

2024-04-30 · Yoonsik Kim, Moonbin Yim, Ka Yeon Song

In this paper, we establish a benchmark for table visual question answering, referred to as the TableVQA-Bench, derived from pre-existing table question-answering (QA) and table structure recognition datasets. It is important to note that existing datasets have not incorporated images or QA pairs, which are two crucial components of TableVQA. As such, the primary objective of this paper is to obtain these necessary components. Specifically, images are sourced either through the application of a \textit{stylesheet} or by employing the proposed table rendering system. QA pairs are generated by exploiting the large language model (LLM) where the input is a text-formatted table. Ultimately, the completed TableVQA-Bench comprises 1,500 QA pairs. We comprehensively compare the performance of various multi-modal large language models (MLLMs) on TableVQA-Bench. GPT-4V achieves the highest accuracy among commercial and open-sourced MLLMs from our experiments. Moreover, we discover that the number of vision queries plays a significant role in TableVQA performance. To further analyze the capabilities of MLLMs in comparison to their LLM backbones, we investigate by presenting image-formatted tables to MLLMs and text-formatted tables to LLMs, respectively. Our findings suggest that processing visual inputs is more challenging than text inputs, as evidenced by the lower performance of MLLMs, despite generally requiring higher computational costs than LLMs. The proposed TableVQA-Bench and evaluation codes are available at \href{https://github.com/naver-ai/tablevqabench}{https://github.com/naver-ai/tablevqabench}.

📄 PDF Abstract BibTeX arXiv:2404.19205

Code (1)

naver-ai/tablevqabench 공식 구현

Tasks

Language ModellingLarge Language ModelQuestion AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

ExpliCIT-QA: Explainable Code-Based Image Table Question Answering

2025-07-15 · Maximiliano Hormazábal Lagos, Álvaro Bueno Sáez, Pedro Alonso Doval, Jorge Alcalde Vesteiro 외 arxiv

We present ExpliCIT-QA, a system that extends our previous MRT approach for tabular question answering into a multimodal pipeline capable of handling complex table images and providing explainable answers. ExpliCIT-QA fo…

Question AnsweringCode Generation

DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates

2026-04-17 · Laziz Hamdi, Amine Tamasna, Thierry Paquet arxiv

Tables condense key transactional and administrative information into compact layouts, but practical extraction requires more than text recognition: systems must also recover structure (rows, columns, merged cells, heade…

Visual Question AnsweringTable Recognition

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

2024-06-27 · Shubhankar Singh, Purvi Chaurasia, Yerram Varun, Pranshu Pandya 외

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities …

Decision MakingLogical ReasoningQuestion AnsweringSpatial Reasoning+2

ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

2022-03-19 · Findings (ACL) 2022 5 · Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty 외

Charts are very popular for analyzing data. When exploring charts, people often ask a variety of complex reasoning questions that involve several logical and arithmetic operations. They also commonly refer to visual feat…

Chart Question AnsweringLogical ReasoningQuestion Answering

MaXM: Towards Multilingual Visual Question Answering

2022-09-12 · Soravit Changpinyo, Linting Xue, Michal Yarom, Ashish V. Thapliyal 외

Visual Question Answering (VQA) has been primarily studied through the lens of the English language. Yet, tackling VQA in other languages in the same manner would require a considerable amount of resources. In this paper…

Question AnsweringTranslationVisual Question AnsweringVisual Question Answering (VQA)