paper-with-me

Papers

The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors

2026-03-01 · Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo arxiv

Effective mathematics education requires identifying and responding to students' mistakes. For AI to support pedagogical applications, models must perform well across different levels of student proficiency. Our work provides an extensive, year-long snapshot of how 11 vision-language models (VLMs) perform on DrawEduMath, a QA benchmark involving real students' handwritten, hand-drawn responses to math problems. We find that models' weaknesses concentrate on a core component of math education: student error. All evaluated VLMs underperform when describing work from students who require more pedagogical help, and across all QA, they struggle the most on questions related to assessing student error. Thus, while VLMs may be optimized to be math problem solving experts, our results suggest that they require alternative development incentives to adequately support educational use cases.

📄 PDF Abstract BibTeX arXiv:2603.00925

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images

2025-01-24 · Sami Baral, Li Lucy, Ryan Knight, Alice Ng 외

In real-world settings, vision language models (VLMs) should robustly handle naturalistic, noisy visual content as well as domain-specific language and concepts. For example, K-12 educators using digital learning platfor…

Math

JEEM: Vision-Language Understanding in Four Arabic Dialects

2025-03-27 · Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali, Abdelrahman Mohamed 외

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. JEEM includes the tasks of image …

Image CaptioningQuestion AnsweringVisual Question Answering

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding

2026-05-19 · Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh, Hung-Ting Su 외 arxiv

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained comprehension crucial for real-world applications requiring nuanced in…

Video Question AnsweringSpatial Reasoning

Cryptocurrency in the Aftermath: Unveiling the Impact of the SVB Collapse

2023-09-15 · Qin Wang, Guangsheng Yu, Shiping Chen

In this paper, we explore the aftermath of the Silicon Valley Bank (SVB) collapse, with a particular focus on its impact on crypto markets. We conduct a multi-dimensional investigation, which includes a factual summary, …

A Zero-Shot Open-Vocabulary Pipeline for Dialogue Understanding

2024-09-24 · Abdulfattah Safa, Gözde Gül Şahin

Dialogue State Tracking (DST) is crucial for understanding user needs and executing appropriate system actions in task-oriented dialogues. Majority of existing DST methods are designed to work within predefined ontologie…

Dialogue State TrackingDialogue Understandingdomain classificationQuestion Answering