paper-with-me

홈 › Papers

IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization

2026-01-21 · Shuai Wang, Yaoming Yang, Bingdong Li, Hao Hao, Aimin Zhou arxiv

Learning Path Recommendation (LPR) aims to generate personalized sequences of learning items that maximize long-term learning effect while respecting pedagogical principles and operational constraints. Although large language models (LLMs) offer rich semantic understanding for free-form recommendation, applying them to long-horizon LPR is challenging due to (i) misalignment with pedagogical objectives such as the Zone of Proximal Development (ZPD) under sparse, delayed feedback, (ii) scarce and costly expert demonstrations, and (iii) multi-objective interactions among learning effect, difficulty scheduling, length controllability, and trajectory diversity. To address these issues, we propose IB-GRPO (Indicator-Based Group Relative Policy Optimization), an indicator-guided alignment approach for LLM-based LPR. To mitigate data scarcity, we construct hybrid expert demonstrations via Genetic Algorithm search and teacher RL agents and warm-start the LLM with supervised fine-tuning. Building on this warm-start, we design a within-session ZPD alignment score for difficulty scheduling. IB-GRPO then uses the $I_{ε+}$ dominance indicator to compute group-relative advantages over multiple objectives, avoiding manual scalarization and improving Pareto trade-offs. Experiments on ASSIST09 and Junyi using the KES simulator with a Qwen2.5-7B backbone show consistent improvements over representative RL and LLM baselines.

📄 PDF Abstract BibTeX arXiv:2601.14686

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Personalized Learning Path Planning with Goal-Driven Learner State Modeling

2025-10-15 · Joy Jia Yin Lim, Ye He, Jifan Yu, Xin Cong 외 arxiv

Personalized Learning Path Planning (PLPP) aims to design adaptive learning paths that align with individual goals. While large language models (LLMs) show potential in personalizing learning experiences, existing approa…

Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach

2025-03-26 · Xuying Li, Zhuo Li, Yuji Kosuga, Victor Bian

Aligning large language models (LLMs) with human values and safety constraints is challenging, especially when objectives like helpfulness, truthfulness, and avoidance of harm conflict. Reinforcement Learning from Human …

Text Generation

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

2026-07-25 · Xiaokun Wang, Siyu Song, Wentao Liu, Xiaodong Zou arxiv

Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedag…

Reinforcement Learning

TreeAdv: Tree-Structured Advantage Redistribution for Group-Based RL

2026-01-07 · Lang Cao, Hui Ruan, Yongqian Li, Peng Chao 외 arxiv

Reinforcement learning with group-based objectives, such as Group Relative Policy Optimization (GRPO), is a common framework for aligning large language models on complex reasoning tasks. However, standard GRPO treats ea…

Reinforcement Learning

REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models

2025-01-04 · Jian Hu

Reinforcement Learning from Human Feedback (RLHF) has emerged as a critical approach for aligning large language models with human preferences, witnessing rapid algorithmic evolution through methods such as Proximal Poli…

Computational Efficiency