paper-with-me

홈 › Papers

Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams

2024-11-07 · Adriana Caraeni, Alexander Scarlatos, Andrew Lan

Recent advances in generative artificial intelligence (AI) have shown promise in accurately grading open-ended student responses. However, few prior works have explored grading handwritten responses due to a lack of data and the challenge of combining visual and textual information. In this work, we leverage state-of-the-art multi-modal AI models, in particular GPT-4o, to automatically grade handwritten responses to college-level math exams. Using real student responses to questions in a probability theory exam, we evaluate GPT-4o's alignment with ground-truth scores from human graders using various prompting techniques. We find that while providing rubrics improves alignment, the model's overall accuracy is still too low for real-world settings, showing there is significant room for growth in this task.

📄 PDF Abstract BibTeX arXiv:2411.05231

Code (0)

등록된 구현이 없습니다.

Tasks

Math

Similar Papers 제목 키워드 기반

Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs

2026-05-18 · Jacob Levine, Miguel Aenlle, Craig Zilles, Matthew West 외 arxiv

Automated grading systems have enabled scalable assessment for many response types, but handwritten mathematics remains a barrier due to the complexity of multi-step solutions. Vision-capable large language models (LLMs)…

Grading Handwritten Engineering Exams with Multimodal Large Language Models

2026-01-02 · Janez Perš, Jon Muhovič, Andrej Košir, Boštjan Murovec arxiv

Handwritten STEM exams capture open-ended reasoning and diagrams, but manual grading is slow and difficult to scale. We present an end-to-end workflow for grading scanned handwritten engineering quizzes with multimodal l…

EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions

2026-01-23 · Weiyu Sun, Liangliang Chen, Yongnuo Cai, Huiru Xie 외 arxiv

Multimodal Large Language Models (MLLMs) hold significant promise for revolutionizing traditional education and reducing teachers' workload. However, accurately interpreting unconstrained STEM student handwritten solutio…

VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions

2025-10-26 · Thu Phuong Nguyen, Duc M. Nguyen, Hyotaek Jeon, Hyunwook Lee 외 arxiv

Automatically assessing handwritten mathematical solutions is an important problem in educational technology with practical applications, but it remains a significant challenge due to the diverse formats, unstructured la…

Reinforcement Learning

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models

2026-06-09 · Hartwig Grabowski arxiv

Correcting handwritten exams by hand is time-consuming and error-prone, particularly for large cohorts, while fully digital exams tend to force a didactic narrowing towards closed question formats. A practical middle gro…