paper-with-me

Papers

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

2026-02-01 · Conrad Borchers, Jill-Jênn Vie, Roger Azevedo arxiv

Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive judgments? Existing evaluations emphasize problem-solving accuracy, overlooking the fragmented and imperfect reasoning that characterizes human learning. We evaluate LLMs as novices using 630 think-aloud utterances from multi-step chemistry tutoring problems with problem-solving logs of student hint use, attempts, and problem context. We compare LLM-generated reasoning to human learner utterances under minimal and extended contextual prompting, and assess the models' ability to predict step-level learner success. Although GPT-4.1 generates fluent and contextually appropriate continuations, its reasoning is systematically over-coherent, verbose, and less variable than human think-alouds. These effects intensify with a richer problem-solving context during prompting. Learner performance was consistently overestimated. These findings highlight epistemic limitations of simulating learning with LLMs. We attribute these limitations to LLM training data, including expert-like solutions devoid of expressions of affect and working memory constraints during problem solving. Our evaluation framework can guide future design of adaptive systems that more faithfully support novice learning and self-regulation using generative artificial intelligence.

📄 PDF Abstract BibTeX arXiv:2602.01015

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Using Think-Aloud Data to Understand Relations between Self-Regulation Cycle Characteristics and Student Performance in Intelligent Tutoring Systems

2023-12-09 · Conrad Borchers, Jiayi Zhang, Ryan S. Baker, Vincent Aleven

Numerous studies demonstrate the importance of self-regulation during learning by problem-solving. Recent work in learning analytics has largely examined students' use of SRL concerning overall learning gains. Limited re…

Designing Conversational AI to Support Think-Aloud Practice in Technical Interview Preparation for CS Students

2025-07-19 · Taufiq Daryanto, Sophia Stil, Xiaohan Ding, Daniel Manesh 외 arxiv

One challenge in technical interviews is the think-aloud process, where candidates verbalize their thought processes while solving coding tasks. Despite its importance, opportunities for structured practice remain limite…

Scaling up the think-aloud method

2025-05-29 · Daniel Wurgaft, Ben Prystawski, Kanishk Gandhi, Cedegao E. Zhang 외

The think-aloud method, where participants voice their thoughts as they solve a task, is a valuable source of rich data about human reasoning processes. Yet, it has declined in popularity in contemporary cognitive scienc…

Mathematical Reasoning

Exploring How Multiple Levels of GPT-Generated Programming Hints Support or Disappoint Novices

2024-04-02 · Ruiwei Xiao, Xinying Hou, John Stamper

Recent studies have integrated large language models (LLMs) into diverse educational contexts, including providing adaptive programming hints, a type of feedback focuses on helping students move forward during problem-so…

A Comparison of the Validity of Measurement Methods for the General English Proficiency through Dictation and Read-Aloud Performances

2021-11-16 · ACL ARR November 2021 11 · Anonymous

This paper compares three measurement methods for the general proficiency of learners of English as a second language (GEP). If students’ GEP can be measured on course materials frequently, for instance, at the beginning…