paper-with-me

홈 › Papers

Auditable Release Control for Pedagogical Leakage in LLM Tutors

2026-08-01 · Nizam Kadir arxiv

Large language model tutors can be correct and helpful yet disclose an answer or decisive reasoning before that disclosure is authorized. We formalize this state- and action-dependent failure as pedagogical leakage and introduce an authorization-aware complete-mediation boundary. A selector emits one of five disclosure contracts, trusted policy gates privileged modes, and a renderer proposes language. A single release function applies inspectable checks, optional cumulative verification, and action-specific fallback; replayable traces separate selection, generation, verification, and enforcement failures. Matched component attribution exposes a safety-utility frontier. On 599 fixed Gemini 3.5 proposals, strict mediation reduces blinded three-model panel-majority leakage flags from 181 to 0 (paired problem-cluster difference -30.22 points, 95% CI [-35.00,-25.72]), while replacing 581 responses and lowering helpfulness. Checker-triggered fallback alone yields 11 majority flags; adding the semantic verifier yields 14 and no reliable marginal gain. A global A1 scaffold yields 0 majority and 54 any-judge flags, outperforming fitted Q on automatic safety and utility. In an externally timestamped replication over 40 unseen problem clusters and 480 attack sequences, high-assurance release reduces majority flags from 42 to 8 (-7.08 points, 95% CI [-13.13,-2.29]); seven failures persist, one is introduced, and mean helpfulness falls by .192. These results establish an auditable release boundary and failure attribution under declared contracts, not universal semantic safety or learning gains.

📄 PDF Abstract BibTeX arXiv:2608.00515

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks

2026-04-20 · Jin Zhao, Marta Knežević, Tanja Käser arxiv

Large Language Models (LLMs) are increasingly used in education, yet their default helpfulness often conflicts with pedagogical principles. Prior work evaluates pedagogical quality via answer leakage-the disclosure of co…

Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors

2024-12-12 · Kaushal Kumar Maurya, KV Aditya Srivatsa, Kseniia Petukhova, Ekaterina Kochmar

In this paper, we investigate whether current state-of-the-art large language models (LLMs) are effective as AI tutors and whether they demonstrate pedagogical abilities necessary for good AI tutoring in educational dial…

Question Answering

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

2026-05-28 · Qikai Chang, Zhenrong Zhang, Linbo Chen, Pengfei Hu 외 arxiv

Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide progressive Socratic guidance and balance multiple pedagogical objectives…

Reinforcement LearningResponse Generation

CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming

2026-07-06 · H. Chad Lane, Bryson Kageler arxiv

Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models. Small language models (SLMs) offer a promising a…

Prompt Engineering

AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms

2025-12-29 · LearnLM Team, Eedi, :, Albert Wang 외 arxiv

One-to-one tutoring is widely considered the gold standard for personalized education, yet it remains prohibitively expensive to scale. To evaluate whether generative AI might help expand access to this resource, we cond…