paper-with-me

Papers

Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors

2025-07-11 · Ekaterina Kochmar, Kaushal Kumar Maurya, Kseniia Petukhova, KV Aditya Srivatsa, Anaïs Tack, Justin Vasselli arxiv

This shared task has aimed to assess pedagogical abilities of AI tutors powered by large language models (LLMs), focusing on evaluating the quality of tutor responses aimed at student's mistake remediation within educational dialogues. The task consisted of five tracks designed to automatically evaluate the AI tutor's performance across key dimensions of mistake identification, precise location of the mistake, providing guidance, and feedback actionability, grounded in learning science principles that define good and effective tutor responses, as well as the track focusing on detection of the tutor identity. The task attracted over 50 international teams across all tracks. The submitted models were evaluated against gold-standard human annotations, and the results, while promising, show that there is still significant room for improvement in this domain: the best results for the four pedagogical ability assessment tracks range between macro F1 scores of 58.34 (for providing guidance) and 71.81 (for mistake identification) on three-class problems, with the best F1 score in the tutor identification track reaching 96.98 on a 9-class task. In this paper, we overview the main findings of the shared task, discuss the approaches taken by the teams, and analyze their performance. All resources associated with this task are made publicly available to support future research in this critical domain.

📄 PDF Abstract BibTeX arXiv:2507.10579

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors

2025-06-12 · Numaan Naeem, Sarfraz Ahmad, Momina Ahsan, Hasan Iqbal

This paper presents our system for Track 1: Mistake Identification in the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors. The task involves evaluating whether a tutor's response correctly ide…

Language ModelingLanguage ModellingLarge Language ModelMathematical Reasoning+1

Small, Private Language Models as Teammates for Educational Assessment Design

2026-05-14 · Chris Davis Jaldi, Anmol Saini, Shan Zhang, Noah Schroeder 외 arxiv

Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bl…

Question Generation

STEM Faculty Perspectives on Generative AI in Higher Education

2026-03-04 · Akila de Silva, Isabel Hyo Jung Song, Hui Yang, Shah Rukh Humayoun arxiv

Generative artificial intelligence (GenAI) tools are increasingly present in higher education, yet their adoption has been largely student-driven, requiring instructors to respond to technologies already embedded in clas…

Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education

2026-05-01 · Juvy C. Grume, John Paul P. Miranda, Aileen P. De Leon, Jordan L. Salenga 외 arxiv

GenAI systems such as ChatGPT are increasingly discussed in programming education, but the ways in which the research literature conceptualizes and frames their role remain unclear. This chapter applies text mining to pu…

A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agents

2025-08-02 · Clayton Cohn, Surya Rayala, Namrata Srivastava, Joyce Horn Fonteles 외 arxiv

Large language models (LLMs) present new opportunities for creating pedagogical agents that engage in meaningful dialogue to support student learning. However, current LLM systems used in classrooms often lack the solid …