paper-with-me

Papers

Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles

2025-12-23 · Ramatu Oiza Abdulsalam, Segun Aroyehun arxiv

Recent work has explored the use of large language models (LLMs) to generate tutoring responses in mathematics, yet it remains unclear how closely their instructional behavior aligns with expert human practice. We analyze a dataset of math remediation dialogues in which expert tutors, novice tutors, and seven LLMs of varying sizes, comprising both open-weight and commercial models, respond to the same student errors. We examine instructional strategies and linguistic characteristics of tutoring responses, including uptake (restating and revoicing), pressing for accuracy and reasoning, lexical diversity, readability, politeness, and agency. We find that expert tutors produce higher-quality responses than novices, and that larger LLMs generally receive higher pedagogical quality ratings than smaller models, approaching expert performance on average. However, LLMs exhibit systematic differences in their instructional profiles: they underuse discursive strategies characteristic of expert tutors while generating longer, more lexically diverse, and more polite responses. Regression analyses show that pressing for accuracy and reasoning, restating and revoicing, and lexical diversity, are positively associated with perceived pedagogical quality, whereas higher levels of agentic and polite language are negatively associated. These findings highlight the importance of analyzing instructional strategies and linguistic characteristics when evaluating tutoring responses across human tutors and intelligent tutoring systems.

📄 PDF Abstract BibTeX arXiv:2512.20780

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors

2025-02-26 · Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur 외

Evaluating the pedagogical capabilities of AI-based tutoring models is critical for making guided progress in the field. Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical ab…

Benchmarking

Towards Pedagogically Aligned LLM Tutors for Math Mistake Remediation

2026-06-19 · Kseniia Petukhova, Tien Dat Nguyen, Ekaterina Kochmar arxiv

Large language models have strong potential for use in intelligent tutoring systems, but they often fail to follow effective pedagogical strategies, such as guiding students without revealing final answers. We study the …

Enabling Multi-Agent Systems as Learning Designers: Applying Learning Sciences to AI Instructional Design

2025-08-20 · Jiayi Wang, Ruiwei Xiao, Xinying Hou, John Stamper arxiv

K-12 educators are increasingly using Large Language Models (LLMs) to create instructional materials. These systems excel at producing fluent, coherent content, but often lack support for high-quality teaching. The reaso…

Prompt Engineering

The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors

2026-03-01 · Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight 외 arxiv

Effective mathematics education requires identifying and responding to students' mistakes. For AI to support pedagogical applications, models must perform well across different levels of student proficiency. Our work pro…

Design and Evaluation for a Prototype of an Online Tool to Access Mathematics Notions in Sign Language

2020-05-01 · LREC 2020 5 · Camille Nadal, Christophe Collet

The Sign{'}Maths project aims at giving access to pedagogical resources in Sign Language (SL). It will provide Deaf students and teachers with mathematics vocabulary in SL, this in order to contribute to the standardisat…

Navigate