paper-with-me

Papers Chart Understanding

“Chart Understanding” 태그가 달린 논문 47편 · 필터 해제

ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering

2025-05-29 · Jingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao 외

Chart question answering (CQA) has become a critical multimodal task for evaluating the reasoning capabilities of vision-language models. While early approaches have shown promising performance by focusing on visual feat…

Chart Question AnsweringChart UnderstandingInstruction FollowingOptical Character Recognition (OCR)+1

InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts

2025-05-25 · Minzhi Lin, Tianchi Xie, Mengchen Liu, Yilin Ye 외

Understanding infographic charts with design-driven visual elements (e.g., pictograms, icons) requires both visual recognition and reasoning, posing challenges for multimodal large language models (MLLMs). However, exist…

Chart UnderstandingQuestion AnsweringVisual Question Answering

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding

2025-05-25 · Muye Huang, Lingling Zhang, Jie Ma, Han Lai 외

Charts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal larg…

Chart UnderstandingLogical Reasoningmultimodal interactionVisual Reasoning

ChartLens: Fine-grained Visual Attribution in Charts

2025-05-25 · Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt 외

The growing capabilities of multimodal large language models (MLLMs) have advanced tasks like chart understanding. However, these models often suffer from hallucinations, where generated text sequences conflict with the …

Chart Understanding

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation

2025-05-24 · Zhen Li, Yukai Guo, Duan Li, Xinyuan Guo 외

Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and structural richness poses challenges for la…

BenchmarkingChart UnderstandingCode GenerationMultimodal Reasoning

OrionBench: A Benchmark for Chart and Human-Recognizable Object Detection in Infographics

2025-05-23 · Jiangning Zhu, Yuxing Zhou, Zheng Wang, Juntao Yao 외

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of …

Chart Understandingobject-detectionObject DetectionVisual Grounding

ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding

2025-05-21 · Yifan Wu, Lutao Yan, Leixian Shen, Yinan Mei 외

The emergence of Multi-modal Large Language Models (MLLMs) presents new opportunities for chart understanding. However, due to the fine-grained nature of these tasks, applying MLLMs typically requires large, high-quality…

Chart Question AnsweringChart UnderstandingQuestion AnsweringRetrieval

ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models

2025-05-19 · Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake 외

Chart understanding presents a unique challenge for large vision-language models (LVLMs), as it requires the integration of sophisticated textual and visual reasoning capabilities. However, current LVLMs exhibit a notabl…

Chart Question AnsweringChart UnderstandingQuestion AnsweringVisual Reasoning

ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing

2025-05-17 · Xuanle Zhao, Xuexin Liu, Haoyue Yang, Xianzhen Luo 외

Although multimodal large language models (MLLMs) show promise in generating chart rendering code, chart editing presents a greater challenge. This difficulty stems from its nature as a labor-intensive task for humans th…

Chart Understanding

ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering

2025-04-07 · Ahmed Masry, Mohammed Saidul Islam, Mahir Ahmed, Aayush Bajaj 외

Charts are ubiquitous, as people often use them to analyze data, answer questions, and discover critical insights. However, performing complex analytical tasks with charts requires significant perceptual and cognitive ef…

Chart Question AnsweringChart UnderstandingMultiple-choiceQuestion Answering

RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning

2025-03-29 · Alexander Vogel, Omar Moured, Yufan Chen, Jiaming Zhang 외

Recently, Vision Language Models (VLMs) have increasingly emphasized document visual grounding to achieve better human-computer interaction, accessibility, and detailed understanding. However, its application to visualiz…

Chart Question AnsweringChart UnderstandingQuestion AnsweringVisual Grounding

On the Perception Bottleneck of VLMs for Chart Understanding

2025-03-24 · Junteng Liu, Weihao Zeng, Xiwen Zhang, Yijun Wang 외

Chart understanding requires models to effectively analyze and reason about numerical data, textual elements, and complex visual components. Our observations reveal that the perception capabilities of existing large visi…

Chart UnderstandingContrastive Learning

Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding

2025-02-17 · Kung-Hsiang Huang, Can Qin, Haoyi Qiu, Philippe Laban 외

Vision Language Models (VLMs) have achieved remarkable progress in multimodal tasks, yet they often struggle with visual arithmetic, seemingly simple capabilities like object counting or length comparison, which are esse…

Arithmetic ReasoningChart UnderstandingDecoderMath+1

ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

2025-01-11 · Xuanle Zhao, Xianzhen Luo, Qi Shi, Chi Chen 외

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks. However, interpreting charts with textual descriptions often leads to information loss, as it fails to full…

Chart UnderstandingCode GenerationLanguage ModelingLanguage Modelling+2

ChartAdapter: Large Vision-Language Model for Chart Summarization

2024-12-30 · Peixin Xu, Yujuan Ding, Wenqi Fan

Chart summarization, which focuses on extracting key information from charts and interpreting it in natural language, is crucial for generating and delivering insights through effective and accessible data analysis. Trad…

Chart Understandingcross-modal alignmentLanguage ModelingLanguage Modelling+1

AskChart: Universal Chart Understanding through Textual Enhancement

2024-12-26 · Xudong Yang, Yifan Wu, Yizhang Zhu, Nan Tang 외

Chart understanding tasks such as ChartQA and Chart-to-Text involve automatically extracting and interpreting key information from charts, enabling users to query or convert visual data into structured formats. State-of-…

Chart UnderstandingMixture-of-Experts

SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images

2024-12-23 · Risa Shinoda, Kuniaki Saito, Shohei Tanaka, Tosho Hirasawa 외

Building a large-scale figure QA dataset requires a considerable amount of work, from gathering and selecting figures to extracting attributes like text, numbers, and colors, and generating QAs. Although recent developme…

Chart Question AnsweringChart Understanding

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

2024-12-13 · Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu 외

We present DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL, through two key major upgrades. For the vision component…

Chart UnderstandingMixture-of-ExpertsOptical Character RecognitionQuestion Answering+3

Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends

2024-10-05 · Mirna Al-Shetairy, Hanan Hindy, Dina Khattab, Mostafa M. Aref

In recent years, interest in vision-language tasks has grown, especially those involving chart interactions. These tasks are inherently multimodal, requiring models to process chart images, accompanying text, underlying …

BenchmarkingChart UnderstandingOptical Character Recognition (OCR)Prompt Engineering+1

SynChart: Synthesizing Charts from Language Models

2024-09-25 · Mengchen Liu, Qixiu Li, Dongdong Chen, Dong Chen 외

With the release of GPT-4V(O), its use in generating pseudo labels for multi-modality tasks has gained significant popularity. However, it is still a secret how to build such advanced models from its base large language …

Chart Understanding
1–20 / 47 다음 →