paper-with-me

홈 › Papers

Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors

2024-07-12 · Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan

Large language models (LLMs) present an opportunity to scale high-quality personalized education to all. A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving. However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors. Inspired by real-world teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation. We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers. We show empirically that finding the mistake in a student solution is challenging for current models. We propose and evaluate several verifiers for detecting these errors. Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines.

📄 PDF Abstract BibTeX arXiv:2407.09136

Code (1)

eth-lre/verify-then-generate 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language ModelMathResponse Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

LLMs cannot spot math errors, even when allowed to peek into the solution

2025-09-01 · KV Aditya Srivatsa, Kaushal Kumar Maurya, Ekaterina Kochmar arxiv

Large language models (LLMs) demonstrate remarkable performance on math word problems, yet they have been shown to struggle with meta-reasoning tasks such as identifying errors in student solutions. In this work, we inve…

Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics

2026-05-31 · Mohammad Amanlou, Yasaman Amou-Jafari, Mehrad Livian, Fatemeh Boloukazari 외 arxiv

Large language models (LLMs) are increasingly entering students' learning practices, but their educational value may depend on whether they are used to support reasoning or to complete tasks without engaging in the under…

Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information

2026-04-17 · Yao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen Liu arxiv

The significant computational demands of large language models have increased interest in distilling reasoning abilities into smaller models via Chain-of-Thought (CoT) distillation. Current CoT distillation methods mainl…

SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning

2023-08-01 · Ning Miao, Yee Whye Teh, Tom Rainforth

The recent progress in large language models (LLMs), especially the invention of chain-of-thought prompting, has made it possible to automatically answer questions by stepwise reasoning. However, when faced with more com…

GSM8KMathQuestion Answering

Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning

2025-12-17 · Jiaqi Xu, Cuiling Lan, Xuejin Chen, Yan Lu arxiv

Human beings solve complex problems through critical thinking, where reasoning and evaluation are intertwined to converge toward correct solutions. However, most existing large language models (LLMs) treat the reasoning …

Reinforcement LearningMathematical Reasoning