paper-with-me

Papers

Towards Reward Modeling for AI Tutors in Math Mistake Remediation

2026-03-25 · Kseniia Petukhova, Ekaterina Kochmar arxiv

Evaluating the pedagogical quality of AI tutors remains challenging: standard NLG metrics do not determine whether responses identify mistakes, scaffold reasoning, or avoid revealing the answers. For the task of mistake remediation, we derive a hierarchy of pedagogical aspects from human pairwise preferences on MRBench, and synthesize minimally contrastive response pairs that differ along key aspects (e.g., mistake identification and location, targetedness, scaffolding, actionability, clarity, and coherence). We develop and release Bradley-Terry preference models trained on weighted-sum rankings that we automatically create from MRBench, synthetic pairs, and data combinations. Using only synthetic data, our best model reaches 0.69 pairwise accuracy on a human preference test, and combining weighted-sum data with targeted synthetic groups improves accuracy to 0.74, outperforming larger general-purpose reward models while using only a 0.5B-parameter backbone.

📄 PDF Abstract BibTeX arXiv:2603.24375

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Pedagogically Aligned LLM Tutors for Math Mistake Remediation

2026-06-19 · Kseniia Petukhova, Tien Dat Nguyen, Ekaterina Kochmar arxiv

Large language models have strong potential for use in intelligent tutoring systems, but they often fail to follow effective pedagogical strategies, such as guiding students without revealing final answers. We study the …

Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes

2023-10-16 · Rose E. Wang, Qingyang Zhang, Carly Robinson, Susanna Loeb 외

Scaling high-quality tutoring remains a major challenge in education. Due to growing demand, many platforms employ novice tutors who, unlike experienced educators, struggle to address student mistakes and thus fail to se…

Decision MakingMath

MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors

2025-05-24 · Baraa Hikal, Mohamed Basem, Islam Oshallah, Ali Hamdi

We present MSA-MathEval, our submission to the BEA 2025 Shared Task on evaluating AI tutor responses across four instructional dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability. …

Language ModelingLanguage ModellingMath

Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors

2025-07-11 · Ekaterina Kochmar, Kaushal Kumar Maurya, Kseniia Petukhova, KV Aditya Srivatsa 외 arxiv

This shared task has aimed to assess pedagogical abilities of AI tutors powered by large language models (LLMs), focusing on evaluating the quality of tutor responses aimed at student's mistake remediation within educati…

Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles

2025-12-23 · Ramatu Oiza Abdulsalam, Segun Aroyehun arxiv

Recent work has explored the use of large language models (LLMs) to generate tutoring responses in mathematics, yet it remains unclear how closely their instructional behavior aligns with expert human practice. We analyz…