paper-with-me

홈 › Papers

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

2026-07-09 · Yi Zhang, Julia Rayz arxiv

We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional intent while shifting its cognitive demand toward specified learning objectives. We apply this framework to programming tasks in computer science education to study the gap between solving tasks and adapting them for learners. Using revised Bloom's Taxonomy as an operational scale of cognitive demand, we evaluate two intervention settings: general difficulty control, where models are asked to make tasks harder or easier, and Bloom's control, where models are asked to target higher or lower Bloom's levels. We evaluate a matched Qwen3-Next model pair, comparing Qwen3-Next-80B-A3B-Instruct with Qwen3-Coder-Next across 2,520 tasks from three benchmarks. The framework reveals a robust directional asymmetry: both models reliably increase cognitive demand, but struggle to lower it. We further characterize these outcomes with semantic-delta clustering and layer-wise Fisher's Discriminant Ratio probing. Within this controlled comparison, the general model shows clearer middle-layer separability for both general difficulty and Bloom-control contrasts, whereas the coder model shows weaker separability for general difficulty and a deeper peak for Bloom-control contrasts. These results show that strong execution performance does not automatically entail Bloom-aligned educational control.

📄 PDF Abstract BibTeX arXiv:2607.08009

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GamED.AI: A Hierarchical Multi-Agent Framework for Automated Educational Game Generation

2026-04-27 · Shiven Agarwal, Yash Shah, Ashish Raj Shekhar, Priyanuj Bordoloi 외 arxiv

We introduce GamEDAI, a hierarchical multi-agent framework that transforms instructor-provided questions into fully playable, pedagogically grounded educational games validated through formal mechanic contracts. Built on…

Spatial Reasoning

Small, Private Language Models as Teammates for Educational Assessment Design

2026-05-14 · Chris Davis Jaldi, Anmol Saini, Shan Zhang, Noah Schroeder 외 arxiv

Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bl…

Question Generation

BloomVQA: Assessing Hierarchical Multi-modal Comprehension

2023-12-20 · Yunye Gong, Robik Shrestha, Jared Claypoole, Michael Cogswell 외

We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus on fact-based memorization and simple re…

Data AugmentationMemorizationMultiple-choiceVisual Question Answering (VQA)

Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment

2025-04-19 · Antoun Yaacoub, Jérôme Da-Rugna, Zainab Assaghir

This study evaluates the integration of Bloom's Taxonomy into OneClickQuiz, an Artificial Intelligence (AI) driven plugin for automating Multiple-Choice Question (MCQ) generation in Moodle. Bloom's Taxonomy provides a st…

ClassificationMultiple-choice

OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models

2026-01-20 · Unggi Lee, Sookbun Lee, Heungsoo Choi, Jinseo Lee 외 arxiv

Large Language Models are increasingly deployed as educational tools, yet existing benchmarks focus on narrow skills and lack grounding in learning sciences. We introduce OpenLearnLM Benchmark, a theory-grounded framewor…