paper-with-me

홈 › Papers

TACLer: Tailored Curriculum Reinforcement Learning for Efficient Reasoning

2026-01-29 · Huiyuan Lai, Malvina Nissim arxiv

Large Language Models (LLMs) have shown remarkable performance on complex reasoning tasks, especially when equipped with long chain-of-thought (CoT) reasoning. However, eliciting long CoT typically requires large-scale reinforcement learning (RL) training, while often leading to overthinking with redundant intermediate steps. To improve learning and reasoning efficiency, while preserving or even enhancing performance, we propose TACLer, a model-tailored curriculum reinforcement learning framework that gradually increases the complexity of the data based on the model's proficiency in multi-stage RL training. TACLer features two core components: (i) tailored curriculum learning that determines what knowledge the model lacks and needs to learn in progressive stages; (ii) a hybrid Thinking/NoThinking reasoning paradigm that balances accuracy and efficiency by enabling or disabling the Thinking mode. Our experiments show that TACLer yields a twofold advantage in learning and reasoning: (i) it reduces computational cost, cutting training compute by over 50% compared to long thinking models and reducing inference token usage by over 42% relative to the base model; and (ii) it improves accuracy by over 9% on the base model, consistently outperforming state-of-the-art Nothinking and Thinking baselines across four math datasets with complex problems.

📄 PDF Abstract BibTeX arXiv:2601.21711

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

2025-05-25 · Shaohao Rui, Kaitao Chen, Weijie Ma, Xiaosong Wang

Recent advances in reinforcement learning with verifiable, rule-based rewards have greatly enhanced the reasoning capabilities and out-of-distribution generalization of VLMs/LLMs, obviating the need for manually crafted …

Out-of-Distribution Generalizationreinforcement-learningReinforcement LearningVisual Question Answering (VQA)+1

Self-Paced Deep Reinforcement Learning

2020-04-24 · NeurIPS 2020 12 · Pascal Klink, Carlo D'Eramo, Jan Peters, Joni Pajarinen

Curriculum reinforcement learning (CRL) improves the learning speed and stability of an agent by exposing it to a tailored series of tasks throughout learning. Despite empirical successes, an open question in CRL is how …

Deep Reinforcement LearningOpen-Ended Question Answeringreinforcement-learningReinforcement Learning+1

Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math

2025-10-30 · Bo Pang, Deqian Kong, Silvio Savarese, Caiming Xiong 외 arxiv

Reinforcement learning (RL) can elicit strong reasoning in large language models (LLMs), yet most open efforts focus on math and code. We propose Reasoning Curriculum, a simple two-stage curriculum that first elicits rea…

Reinforcement Learning

IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning

2026-04-27 · Navya Gupta, Rishitej Reddy Vyalla, Avinash Anand, Chhavi Kirtani 외 arxiv

Curriculum learning helps language models tackle complex reasoning by gradually increasing task difficulty. However, it often fails to generate consistent step-by-step reasoning, especially in multilingual and low-resour…

Reinforcement LearningCross-Lingual TransferMathematical Reasoning

Boosted Curriculum Reinforcement Learning

2021-09-29 · ICLR 2022 4 · Pascal Klink, Carlo D'Eramo, Jan Peters, Joni Pajarinen

Curriculum value-based reinforcement learning (RL) solves a complex target task by reusing action-values across a tailored sequence of related tasks of increasing difficulty. However, finding an exact way of reusing acti…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)