paper-with-me

홈 › Papers

Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning

2026-04-20 · Moiz Imran, Sahan Bulathwela arxiv

Intelligent tutoring systems increasingly provide automated feedback on student work, but robust feedback requires assessing reasoning, not only final answers. We study a failure mode we call the correct answer trap (CAT): models under-detect misconceptions when students reach a correct answer via flawed reasoning. Analysing real student responses from the Eedi mathematics platform, we show that 71% of these failures concentrate in just two question types, both sharing a common structure where flawed reasoning happens to produce the correct numerical answer. Comparing a fine-tuned T5 with a frontier large language model, we find that improved capabilities reduce but do not eliminate the problem (84% vs 57% detection accuracy). Even the best-performing model generates roughly four false alarms for every genuine detection, making stand-alone screening impractical at realistic class sizes. Our findings demonstrate that high overall accuracy can mask critical failures in reasoning assessment, and that careful analysis of student reasoning still benefits from human judgment.

📄 PDF Abstract BibTeX arXiv:2605.23925

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Correct Answer Trap: Pedagogically-Grounded Detection and Feedback for Hidden Misconceptions

2026-06-22 · Moiz Imran, Sahan Bulathwela arxiv

Automated feedback systems that rely on answer correctness will reinforce, rather than address, misconceptions when students reach the correct answer through flawed reasoning. We investigate automatic detection of these …

Novice Learner and Expert Tutor: Evaluating Math Reasoning Abilities of Large Language Models with Misconceptions

2023-10-03 · Naiming Liu, Shashank Sonkar, Zichao Wang, Simon Woodhead 외

We propose novel evaluations for mathematical reasoning capabilities of Large Language Models (LLMs) based on mathematical misconceptions. Our primary approach is to simulate LLMs as a novice learner and an expert tutor,…

MathMathematical ReasoningMisconceptions

Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing

2026-07-06 · Bonan Shen, Dingyan Shang, Youting Wang, Tao Ning arxiv

Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does not guarantee that the answer was derived from the student-facing problem. In real…

Few-shot Question Generation for Personalized Feedback in Intelligent Tutoring Systems

2022-06-08 · Devang Kulshreshtha, Muhammad Shayan, Robert Belfer, Siva Reddy 외

Existing work on generating hints in Intelligent Tutoring Systems (ITS) focuses mostly on manual and non-personalized feedback. In this work, we explore automatically generated questions as personalized feedback in an IT…

Generative Question AnsweringQuestion AnsweringQuestion GenerationQuestion-Generation+2

Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks

2026-04-20 · Jin Zhao, Marta Knežević, Tanja Käser arxiv

Large Language Models (LLMs) are increasingly used in education, yet their default helpfulness often conflicts with pedagogical principles. Prior work evaluates pedagogical quality via answer leakage-the disclosure of co…