paper-with-me

Papers

MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring

2025-10-27 · Tengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su, Yuanyang Liu, Zhihan Zhang, Meng Jiang arxiv

Effective math tutoring requires not only solving problems but also diagnosing students' difficulties and guiding them step by step. While multimodal large language models (MLLMs) show promise, existing benchmarks largely overlook these tutoring skills. We introduce MMTutorBench, the first benchmark for AI math tutoring, consisting of 685 problems built around pedagogically significant key-steps. Each problem is paired with problem-specific rubrics that enable fine-grained evaluation across six dimensions, and structured into three tasks-Insight Discovery, Operation Formulation, and Operation Execution. We evaluate 12 leading MLLMs and find clear performance gaps between proprietary and open-source systems, substantial room compared to human tutors, and consistent trends across input variants: OCR pipelines degrade tutoring quality, few-shot prompting yields limited gains, and our rubric-based LLM-as-a-Judge proves highly reliable. These results highlight both the difficulty and diagnostic value of MMTutorBench for advancing AI tutoring.

📄 PDF Abstract BibTeX arXiv:2510.23477

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset

2026-02-18 · Zhuqian Zhou, Kirk Vanacore, Bakhtawar Ahtisham, Jinsook Lee 외 arxiv

Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrier. In mathematics tutoring transcripts, numeric expressions frequently r…

GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns

2026-05-08 · Sankalan Pal Chowdhury, Junling Wang, Donya Rooein, April Yi Wang 외 arxiv

Several educational domains rely heavily on diagrams and visual cues, yet most existing tutoring datasets are limited to text-only interactions. This limits the development of AI tutors that can teach in visually grounde…

Visual Reasoning

MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors

2025-02-26 · Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur 외

Evaluating the pedagogical capabilities of AI-based tutoring models is critical for making guided progress in the field. Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical ab…

Benchmarking

Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving

2025-02-12 · Steven-Shine Chen, JiMin Lee, Paul Pu Liang

Humans have long relied on visual aids like sketches and diagrams to support reasoning and problem-solving. Visual tools, like auxiliary lines in geometry or graphs in calculus, are essential for understanding complex id…

Mathmultimodal interaction

A Dialogue-Based Framework for Correcting Multimodal Errors in AI-Assisted STEM Education

2026-05-05 · Akshay Syal, Lawrence Swaminathan Xavier Prince, Evin Gultepe, Nik Bear Brown 외 arxiv

Large Language Models (LLMs) are democratizing access to personalized tutoring; however, their effectiveness is hindered by challenges in processing multimodal content, which limits AI's potential to provide equitable, h…