paper-with-me

Papers

InfographicVQA

2021-04-26 · Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito, Dimosthenis Karatzas, Ernest Valveny, C. V Jawahar

Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic understanding of infographic images by using Visual Question Answering technique.To this end, we present InfographicVQA, a new dataset that comprises a diverse collection of infographics along with natural language questions and answers annotations. The collected questions require methods to jointly reason over the document layout, textual content, graphical elements, and data visualizations. We curate the dataset with emphasis on questions that require elementary reasoning and basic arithmetic skills. Finally, we evaluate two strong baselines based on state of the art multi-modal VQA models, and establish baseline performance for the new task. The dataset, code and leaderboard will be made available at http://docvqa.org

📄 PDF Abstract BibTeX arXiv:2104.12756

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics

2025-12-13 · Tue-Thu Van-Dinh, Hoang-Duy Tran, Truong-Binh Duong, Mai-Hanh Pham 외 arxiv

Infographic Visual Question Answering (InfographicVQA) evaluates a model's ability to read and reason over data-rich, layout-heavy visuals that combine text, charts, icons, and design elements. Compared with scene-text o…

Visual Question Answering

Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding

2026-03-24 · Mincheol Kwon, Minseung Lee, Seonga Choi, Miso Choi 외 arxiv

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and inform…

ScreenAI: A Vision-Language Model for UI and Infographics Understanding

2024-02-07 · Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach 외

Screen user interfaces (UIs) and infographics, sharing similar visual language and design principles, play important roles in human communication and human-machine interaction. We introduce ScreenAI, a vision-language mo…

Chart Question AnsweringLanguage ModelingLanguage ModellingQuestion Answering+1

Enhancing Document VQA Models via Retrieval-Augmented Generation

2025-08-26 · Eric López, Artemis Llabrés, Ernest Valveny arxiv

Document Visual Question Answering (Document VQA) must cope with documents that span dozens of pages, yet leading systems still concatenate every page or rely on very large vision-language models, both of which are memor…

Visual Question Answering

Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation

2025-10-23 · Yuhan Liu, Lianhui Qin, Shengjie Wang arxiv

Large Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet they struggle when reasoning over information-intensive images that densely interleave textual annotations with fine-…

Visual Question AnsweringVisual Reasoning