paper-with-me

홈 › Papers

Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math

2026-03-26 · Dingjie Song, Tianlong Xu, Yi-Fan Zhang, Hang Li, Zhiling Yan, Xing Fan, Haoyang Li, Lichao Sun, Qingsong Wen arxiv

Assessing student handwritten scratchwork is crucial for personalized educational feedback but presents unique challenges due to diverse handwriting, complex layouts, and varied problem-solving approaches. Existing educational NLP primarily focuses on textual responses and neglects the complexity and multimodality inherent in authentic handwritten scratchwork. Current multimodal large language models (MLLMs) excel at visual reasoning but typically adopt an "examinee perspective", prioritizing generating correct answers rather than diagnosing student errors. To bridge these gaps, we introduce ScratchMath, a novel benchmark specifically designed for explaining and classifying errors in authentic handwritten mathematics scratchwork. Our dataset comprises 1,720 mathematics samples from Chinese primary and middle school students, supporting two key tasks: Error Cause Explanation (ECE) and Error Cause Classification (ECC), with seven defined error types. The dataset is meticulously annotated through rigorous human-machine collaborative approaches involving multiple stages of expert labeling, review, and verification. We systematically evaluate 16 leading MLLMs on ScratchMath, revealing significant performance gaps relative to human experts, especially in visual recognition and logical reasoning. Proprietary models notably outperform open-source models, with large reasoning models showing strong potential for error explanation. All evaluation data and frameworks are publicly available to facilitate further research.

📄 PDF Abstract BibTeX arXiv:2603.24961

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

Merlin:Empowering Multimodal LLMs with Foresight Minds

2023-11-30 · En Yu, Liang Zhao, Yana Wei, Jinrong Yang 외

Humans possess the remarkable ability to foresee the future to a certain extent based on present observations, a skill we term as foresight minds. However, this capability remains largely under explored within existing M…

Visual Question Answering

MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias

2026-06-16 · Xingming Li, Ao Cheng, Qiyao Sun, Xixiang He 외 arxiv

When vision contradicts text, multimodal large language models (MLLMs) consistently favor text, even when images provide clear evidence otherwise. This bias poses risks for applications requiring visual grounding, yet it…

Visual Grounding

"Mistakes Help Us Grow": Facilitating and Evaluating Growth Mindset Supportive Language in Classrooms

2023-10-16 · Kunal Handa, Margaret Clapper, Jessica Boyle, Rose E Wang 외

Teachers' growth mindset supportive language (GMSL)--rhetoric emphasizing that one's skills can be improved over time--has been shown to significantly reduce disparities in academic achievement and enhance students' lear…

Building Flexible, Scalable, and Machine Learning-ready Multimodal Oncology Datasets

2023-09-30 · Aakash Tripathi, Asim Waqas, Kavya Venkatesan, Yasin Yilmaz 외

The advancements in data acquisition, storage, and processing techniques have resulted in the rapid growth of heterogeneous medical data. Integrating radiological scans, histopathology images, and molecular information w…

Data IntegrationDiagnostic

Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition

2025-08-12 · Mustafa Akben, Vinayaka Gude, Haya Ajjan arxiv

The ability to discern subtle emotional cues is fundamental to human social intelligence. As artificial intelligence (AI) becomes increasingly common, AI's ability to recognize and respond to human emotions is crucial fo…

Emotional IntelligenceEmotion Recognition