paper-with-me

홈 › Papers

Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold

2026-07-30 · Songshuo Lu, Zhi Chen, Yaohua Tang arxiv

A single reinforcement-learning run can produce a strong reasoner yet an incomplete teacher: it often amplifies only a subset of the valid solution modes. We argue that reinforcement learning (RL)-trained policies should therefore be viewed as local probes of a multi-basin reasoning solution manifold, rather than as globally reliable supervisors. Based on this view, we propose an expand-then-compress framework that couples teacher construction with multi-teacher policy distillation. In the expansion stage, Residual Group Relative Policy Optimization (RGRPO) trains a sequence of teachers from a common initialization and redirects each later round toward examples not yet covered by the accumulated teacher union. In the compression stage, reliability-gated Teacher-Union On-policy Distillation (TU-OPD) lets the student learn from its own response prefixes. For each example, only reliable teachers contribute, and their sampled-token OPD losses are weighted by their per-example quality. We further introduce Consensus-Residual Decomposition, which preserves a winner teacher's excess token preferences over its reliable peers, preventing specialist behavior from being suppressed during teacher aggregation. Experiments on mathematical reasoning, code generation, and instruction following show that the resulting Qwen3-1.7B student consistently outperforms the strongest individual teacher across all three domains, yielding relative improvements of 2.0%, 8.3%, and 6.9%, respectively, while retaining single-model inference. These results establish a simple but powerful principle: stronger students can be obtained not by selecting a single better teacher, but by deliberately constructing and compressing a complementary teacher union.

📄 PDF Abstract BibTeX arXiv:2607.27770

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical ReasoningInstruction FollowingCode Generation

Similar Papers 제목 키워드 기반

Merge-of-Thought Distillation

2025-09-10 · Zhanming Shen, Zeyu Qin, Zenan Huang, Hao Chen 외 arxiv

Efficient reasoning distillation for long chain-of-thought (CoT) models is increasingly constrained by the assumption of a single oracle teacher, despite the practical availability of multiple candidate teachers and grow…

Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions

2025-04-20 · Luyang Fang, Xiaowei Yu, Jiazhang Cai, Yongkai Chen 외

The exponential growth of Large Language Models (LLMs) continues to highlight the need for efficient strategies to meet ever-expanding computational and data demands. This survey provides a comprehensive analysis of two …

Dataset DistillationDiversityKnowledge DistillationSurvey

Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text

2026-05-27 · Tianyang Zhou, Wenbo Chen, Pierre Jinghong Liang, Leman Akoglu arxiv

LLMs have advanced text classification, yet existing paradigms face a trade-off: supervised (label only) fine-tuning is scalable but offers limited reasoning on complex text and lacks broader model transparency, while di…

Reinforcement LearningText Classification

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning

2026-05-12 · Zihao Han, Tiangang Zhang, Huaibin Wang, Yilun Sun arxiv

On-policy self-distillation has become a strong recipe for LLM reasoning, where a privileged teacher supervises the student's own rollouts while conditioning on the reference solution. A design choice shared by nearly al…

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

2026-01-26 · Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang 외 arxiv

Knowledge distillation improves large language model (LLM) reasoning by compressing the knowledge of a teacher LLM to train smaller LLMs. On-policy distillation advances this approach by having the student sample its own…

Knowledge DistillationMathematical ReasoningReinforcement Learning