paper-with-me

홈 › Papers

mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning

2024-04-02 · Jingxuan Wei, Nan Xu, Guiyong Chang, Yin Luo, Bihui Yu, Ruifeng Guo

In the fields of computer vision and natural language processing, multimodal chart question-answering, especially involving color, structure, and textless charts, poses significant challenges. Traditional methods, which typically involve either direct multimodal processing or a table-to-text conversion followed by language model analysis, have limitations in effectively handling these complex scenarios. This paper introduces a novel multimodal chart question-answering model, specifically designed to address these intricate tasks. Our model integrates visual and linguistic processing, overcoming the constraints of existing methods. We adopt a dual-phase training approach: the initial phase focuses on aligning image and text representations, while the subsequent phase concentrates on optimizing the model's interpretative and analytical abilities in chart-related queries. This approach has demonstrated superior performance on multiple public datasets, particularly in handling color, structure, and textless chart questions, indicating its effectiveness in complex multimodal tasks.

📄 PDF Abstract BibTeX arXiv:2404.01548

Code (0)

등록된 구현이 없습니다.

Tasks

Chart Question AnsweringLanguage ModelingLanguage ModellingQuestion Answering

Similar Papers 제목 키워드 기반

CHARTOM: A Visual Theory-of-Mind Benchmark for Multimodal Large Language Models

2024-08-26 · Shubham Bharti, Shiyun Cheng, Jihyun Rho, Jianrui Zhang 외

We introduce CHARTOM, a visual theory-of-mind benchmark for multimodal large language models. CHARTOM consists of specially designed data visualizing charts. Given a chart, a language model needs to not only correctly co…

Language ModelingLanguage Modelling

InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts

2025-05-25 · Minzhi Lin, Tianchi Xie, Mengchen Liu, Yilin Ye 외

Understanding infographic charts with design-driven visual elements (e.g., pictograms, icons) requires both visual recognition and reasoning, posing challenges for multimodal large language models (MLLMs). However, exist…

Chart UnderstandingQuestion AnsweringVisual Question Answering

ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering

2025-05-29 · Jingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao 외

Chart question answering (CQA) has become a critical multimodal task for evaluating the reasoning capabilities of vision-language models. While early approaches have shown promising performance by focusing on visual feat…

Chart Question AnsweringChart UnderstandingInstruction FollowingOptical Character Recognition (OCR)+1

Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering

2025-03-23 · Zixin Chen, Sicheng Song, Kashun Shum, Yanna Lin 외

Misleading chart visualizations, which intentionally manipulate data representations to support specific claims, can distort perceptions and lead to incorrect conclusions. Despite decades of research, misleading visualiz…

BenchmarkingChart Question AnsweringMultiple-choiceQuestion Answering

Benchmarking Multimodal RAG through a Chart-based Document Question-Answering Generation Framework

2025-02-20 · Yuming Yang, Jiang Zhong, Li Jin, Jingwang Huang 외

Multimodal Retrieval-Augmented Generation (MRAG) enhances reasoning capabilities by integrating external knowledge. However, existing benchmarks primarily focus on simple image-text interactions, overlooking complex visu…

BenchmarkingQuestion AnsweringRAGRetrieval+1