paper-with-me

홈 › Papers

"Did my figure do justice to the answer?" : Towards Multimodal Short Answer Grading with Feedback (MMSAF)

2024-12-27 · Pritam Sil, Pushpak Bhattacharyya

Assessments play a vital role in a student's learning process. This is because they provide valuable feedback crucial to a student's growth. Such assessments contain questions with open-ended responses, which are difficult to grade at scale. These responses often require students to express their understanding through textual and visual elements together as a unit. In order to develop scalable assessment tools for such questions, one needs multimodal LLMs having strong comparative reasoning capabilities across multiple modalities. Thus, to facilitate research in this area, we propose the Multimodal Short Answer grading with Feedback (MMSAF) problem along with a dataset of 2,197 data points. Additionally, we provide an automated framework for generating such datasets. As per our evaluations, existing Multimodal Large Language Models (MLLMs) could predict whether an answer is correct, incorrect or partially correct with an accuracy of 55%. Similarly, they could predict whether the image provided in the student's answer is relevant or not with an accuracy of 75%. As per human experts, Pixtral was more aligned towards human judgement and values for biology and ChatGPT for physics and chemistry and achieved a score of 4 or more out of 5 in most parameters.

📄 PDF Abstract BibTeX arXiv:2412.19755

Code (0)

등록된 구현이 없습니다.

Tasks

automatic short answer grading

Similar Papers 제목 키워드 기반

LEAF-QA: Locate, Encode & Attend for Figure Question Answering

2019-07-30 · Ritwick Chaudhry, Sumit Shekhar, Utkarsh Gupta, Pranav Maneriker 외

We introduce LEAF-QA, a comprehensive dataset of $250,000$ densely annotated figures/charts, constructed from real-world open data sources, along with ~2 million question-answer (QA) pairs querying the structure and sema…

Chart Question AnsweringQuestion AnsweringVisual Question Answering (VQA)

BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence

2026-03-09 · Biao Xiang, Soyeon Caren Han, Yihao Ding arxiv

Multi-hop question answering (QA) is widely used to evaluate the reasoning capabilities of large language models, yet most benchmarks focus on final answer correctness and overlook intermediate reasoning, especially in l…

Multi-hop Question Answering

MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models

2026-03-12 · Michiko Yoshitake, Yuta Suzuki, Ryo Igarashi, Yoshitaka Ushiku 외 arxiv

We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems that require accurate interpretation of fi…

Multimodal ReasoningVisual Reasoning

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

2026-06-08 · Muhammad Umer Sheikh, Hassan Abid, Khawar Shehzad, Ufaq Khan 외 arxiv

Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of …

Question Answering

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents

2026-04-07 · Xuan Dong, Huanyang Zheng, Tianhao Niu, Zhe Han 외 arxiv

Scientific research follows multi-turn, multi-step workflows that require proactively searching the literature, consulting figures and tables, and integrating evidence across papers to align experimental settings and sup…