paper-with-me

홈 › Papers

Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training

2025-11-10 · Dake Bu, Wei Huang, Andi Han, Atsushi Nitanda, Hau-San Wong, Qingfu Zhang, Taiji Suzuki arxiv

Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a principled understanding of their effectiveness and limitations remains incomplete. To bridge this gap, we develop an abstract theoretical framework and identify sufficient conditions under which curriculum post-training yields exponential improvements in sample complexity. To substantiate this framework, we model the base model's Chain-of-Thought generation as a state-conditioned autoregressive reasoning tree, and formalize curriculum subtasks as either depth-increasing curricula that progressively extend reasoning horizons or hint-decreasing curricula that gradually remove partial hints. Our analysis shows that reinforcement learning finetuning with both curriculum strategies achieves high accuracy with polynomial sample complexity, whereas non-curriculum counterpart encounters an exponential complexity bottleneck. We further establish analogous guarantees for test-time scaling. Empirical simulations support our theoretical findings. Code is available at https://github.com/DakeBU/Curriculum-Post-training.

📄 PDF Abstract BibTeX arXiv:2511.07372

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

2026-03-18 · Nived Rajaraman, Audrey Huang, Miro Dudik, Robert Schapire 외 arxiv

Chain-of-thought reasoning, where language models expend additional computation by producing thinking tokens prior to final responses, has driven significant advances in model capabilities. However, training these reason…

Reinforcement Learning

Progressive distillation induces an implicit curriculum

2024-10-07 · Abhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Andrej Risteski 외

Knowledge distillation leverages a teacher model to improve the training of a student model. A persistent challenge is that a better teacher does not always yield a better student, to which a common mitigation is to use …

Knowledge Distillation

A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning

2026-05-13 · Jason Gaitonde, Frederic Koehler, Elchanan Mossel, Joonhyung Shin 외 arxiv

We introduce a family of synthetic languages with hierarchical structure -- generated by a broadcast process on trees -- for which the role of context length and reasoning in autoregressive generation can be analyzed pre…

TransCurriculum: Multi-Dimensional Curriculum Learning for Fast & Stable Locomotion

2026-03-14 · Prakhar Mishra, Amir Hossain Raj, Xuesu Xiao, Dinesh Manocha arxiv

High-speed legged locomotion struggles with stability and transfer losses at higher command velocities during deployment. One reason is that most curricula vary difficulty along single axis, for example increase the rang…

Mathematical Reasoning via Self-supervised Skip-tree Training

2020-06-08 · ICLR 2021 1 · Markus N. Rabe, Dennis Lee, Kshitij Bansal, Christian Szegedy

We examine whether self-supervised language modeling applied to mathematical formulas enables logical reasoning. We suggest several logical reasoning tasks that can be used to evaluate language models trained on formal m…

Language ModelingLanguage ModellingLogical ReasoningMathematical Reasoning