paper-with-me

홈 › Papers

Efficient RL for optimizing conversation level outcomes with an LLM-based tutor

2025-07-22 · Hyunji Nam, Omer Gottesman, Amy Zhang, Dean Foster, Emma Brunskill, Lyle Ungar arxiv

Large language models (LLMs) built on existing reinforcement learning with human feedback (RLHF) frameworks typically optimize responses based on immediate turn-level human preferences. However, this approach falls short in multi-turn dialogue settings, such as online math tutoring. We propose a method to enhance LLM-based tutors by representing the dialogue history with a lower-dimensional latent state representation of a student and optimizing a long-term policy to determine high-level actions based on the latent state. The goal is to better align the tutor's behavior with the long-term objective of guiding the student towards solving a target math problem on their own. Our model is lightweight, requiring less computational resources than prior work of training the tutor policy end-to-end to directly output the tutor's next utterance. Our experiment results demonstrate that these modifications lead to improved long-term outcomes compared to prompting in LLM-simulated tutoring tasks.

📄 PDF Abstract BibTeX arXiv:2507.16252

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Ruffle&Riley: Towards the Automated Induction of Conversational Tutoring Systems

2023-09-26 · Robin Schmucker, Meng Xia, Amos Azaria, Tom Mitchell

Conversational tutoring systems (CTSs) offer learning experiences driven by natural language interaction. They are known to promote high levels of cognitive engagement and benefit learning outcomes, particularly in reaso…

Combining Large Language Models with Tutoring System Intelligence: A Case Study in Caregiver Homework Support

2024-12-16 · Devika Venugopalan, Ziwen Yan, Conrad Borchers, Jionghao Lin 외

Caregivers (i.e., parents and members of a child's caring community) are underappreciated stakeholders in learning analytics. Although caregiver involvement can enhance student academic outcomes, many obstacles hinder in…

Large Language ModelMathPrompt Engineering

Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents

2026-06-01 · Aitor Arronte Alvarez, Naiyi Xie Fincham arxiv

Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly used in these systems to provide scalable, personalized feedback. Howe…

Bias Detection

Curiosity as Linguistic Intervention: Using LLM Tutoring Dialogues to Influence Exploratory Learning Behavior

2026-06-21 · Gevindu Ganganath, Pasindu Bolonghege, Qianru Lyu, Pradeep Varakantham 외 arxiv

Large Language Models (LLMs) provide a new opportunity to study how language shapes exploratory cognition because conversational strategies can be systematically manipulated at inference time. We introduce CURIOBOT, a fr…

PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors

2026-01-13 · Donya Rooein, Sankalan Pal Chowdhury, Mariia Eremeeva, Yuan Qin 외 arxiv

Recent advances in large language models (LLMs) demonstrate their potential as educational tutors. However, different tutoring strategies benefit different student personalities, and mismatches can be counterproductive t…