paper-with-me

홈 › Papers

Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning

2025-09-20 · Janak Kapuriya, Anwar Shaikh, Arnav Goel, Medha Hira, Apoorv Singh, Jay Saraf, Sanjana, Vaibhav Nauriyal, Avinash Anand, Zhengkui Wang, Rajiv Ratn Shah arxiv

In this study, we introduce Vision-Caption aware Supervised FineTuning (VCASFT), a novel learning paradigm designed to enhance the performance of smaller Vision Language Models(VLMs) on scientific visual question answering(VQA) tasks. VCASFT leverages image captions as zero-shot prompts alongside question-answer pairs and instruction-tunes models to yield significant performance improvements. To comprehensively evaluate VCASFT, we benchmark it on ScienceQA, which consists of questions across diverse languages, subjects, and fields, demonstrating its adaptability and effectiveness in a variety of educational contexts. Additionally, to further demonstrate the effectiveness of this technique on lowresource languages, we developed HiSciVQA, a dataset comprising 2,245 high-quality, hand-annotated Hindi multimodal Q&A pairs. This dataset addresses the critical need for low-resource language Q&A datasets and serves as a foundation for testing VCASFT. Additionally, we introduce a novel LLM-based evaluation scheme to evaluate VLMs on HiSciVQA which offers deeper insights into model effectiveness surpassing traditional n-gram matching accuracy metrics. We are committed to advancing the field by open-sourcing all code files and the HiSciVQA dataset for the research community.

📄 PDF Abstract BibTeX arXiv:2509.16628

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

2025-07-08 · Prahitha Movva, Naga Harshita Marupaka

Technical reports and articles often contain valuable information in the form of semi-structured data like charts, and figures. Interpreting these and using the information from them is essential for downstream tasks suc…

ArticlesMultimodal ReasoningQuestion AnsweringVisual Question Answering

SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning

2025-11-19 · Wenhan Yu, Zhaoxi Zhang, Wang Chen, Guanqiang Qi 외 arxiv

Scientific documents contain complex multimodal structures, which makes evidence localization and scientific reasoning in Document Visual Question Answering particularly challenging. However, most existing benchmarks eva…

Visual Question Answering

Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

2024-03-01 · Lei LI, Yuqi Wang, Runxin Xu, Peiyi Wang 외

Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes. However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains l…

BenchmarkingMathematical ReasoningQuestion Answering

VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering

2025-11-25 · Yuyi Li, Daoyuan Chen, Zhen Wang, Yutong Lu 외 arxiv

Large Vision-Language Models (LVLMs) show promise for scientific applications, yet open-source models still struggle with Scientific Visual Question Answering (SVQA), namely answering questions about figures from scienti…

Visual Question Answering

Language bias in Visual Question Answering: A Survey and Taxonomy

2021-11-16 · Desen Yuan

Visual question answering (VQA) is a challenging task, which has attracted more and more attention in the field of computer vision and natural language processing. However, the current visual question answering has the p…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)