paper-with-me

홈 › Papers

Multimodal LLMs for Visualization Reconstruction and Understanding

2025-06-26 · Can Liu, Chunlin Da, Xiaoxiao Long, Yuxiao Yang, Yu Zhang, Yong Wang

Visualizations are crucial for data communication, yet understanding them requires comprehension of both visual elements and their underlying data relationships. Current multimodal large models, while effective in natural image understanding, struggle with visualization due to their inability to decode the data-to-visual mapping rules and extract structured information. To address these challenges, we present a novel dataset and train multimodal visualization LLMs specifically designed for understanding. Our approach combines chart images with their corresponding vectorized representations, encoding schemes, and data features. The proposed vector format enables compact and accurate reconstruction of visualization content. Experimental results demonstrate significant improvements in both data extraction accuracy and chart reconstruction quality.

📄 PDF Abstract BibTeX arXiv:2506.21319

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visualization Literacy of Multimodal Large Language Models: A Comparative Study

2024-06-24 · Zhimin Li, Haichao Miao, Valerio Pascucci, Shusen Liu

The recent introduction of multimodal large language models (MLLMs) combine the inherent power of large language models (LLMs) with the renewed capabilities to reason about the multimodal context. The potential usage sce…

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

2026-07-16 · Patrick Phuoc Do, Chau M. Ta, Chaoli Wang arxiv

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (…

See or Recall: A Sanity Check for the Role of Vision in Solving Visualization Question Answer Tasks with Multimodal LLMs

2025-04-14 · Zhimin Li, Haichao Miao, Xinyuan Yan, Valerio Pascucci 외

Recent developments in multimodal large language models (MLLM) have equipped language models to reason about vision and language jointly. This permits MLLMs to both perceive and answer questions about data visualization …

Data VisualizationQuestion Answering

Exploring Multimodal Prompt for Visualization Authoring with Large Language Models

2025-04-18 · Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang 외

Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language…

SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation

2025-03-09 · Zisheng Chen, Chunwei Wang, Xiuwei Chen, Hang Xu 외

We present SemHiTok, a unified image Tokenizer via Semantic-Guided Hierarchical codebook that provides consistent discrete feature representations for multimodal understanding and generation tasks. Recently, unified mult…