paper-with-me

홈 › Papers

Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective

2026-04-10 · Tokio Kajitsuka, Ukyo Honda, Sho Takase arxiv

Chain-of-thought (CoT) distillation transfers reasoning behaviors from a strong teacher to a smaller student, but prior work reports a capacity gap: distillation may fail when the teacher-student capability mismatch is large. We revisit the capacity gap from a practical perspective by re-examining commonly used experimental settings. Notably, we find that CoT distillation often degrades performance compared to the student's pre-distillation baseline, an issue obscured when only post-distillation comparisons are reported. We therefore propose a more realistic evaluation protocol and find that the impact of capacity gap effects does not consistently dominate across tasks and settings, especially when candidate teachers differ substantially in performance. Our results offer practical guidance for selecting teacher-student pairs in CoT distillation.

📄 PDF Abstract BibTeX arXiv:2604.08880

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning

2025-05-20 · Shangziqi Zhao, Jiahao Yuan, Guisong Yang, Usman Naseem

Long chain-of-thought (Long-CoT) reasoning improves accuracy in LLMs, yet its verbose, self-reflective style often hinders effective distillation into small language models (SLMs). We revisit Long-CoT compression through…

Large Language ModelMathematical Reasoning

OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models

2026-05-12 · Yuanhao Yue, Chengyu Wang, Yuanjie Lyu, Lei Shen 외 arxiv

Recent multimodal large language models (MLLMs) have shown strong chain-of-thought (CoT) reasoning ability on vision-language tasks, but their direct deployment in real-world systems is often limited by latency and resou…

Multimodal Reasoning

Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step

2023-06-24 · Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren 외

Chain-of-thought prompting (e.g., "Let's think step-by-step") primes large language models to verbalize rationalization for their predictions. While chain-of-thought can lead to dramatic performance gains, benefits appea…

Diversity

Revisiting Parallel Context Windows: A Frustratingly Simple Alternative and Chain-of-Thought Deterioration

2023-05-24 · Kejuan Yang, Xiao Liu, Kaiwen Men, Aohan Zeng 외

We identify two crucial limitations in the evaluation of recent parallel-integrated method Parallel Context Windows (PCW), which extends the maximum context lengths of language models, e.g., 2048 for LLaMA, by harnessing…

Long-Context Understanding

Small Models Struggle to Learn from Strong Reasoners

2025-02-17 · Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang 외

Large language models (LLMs) excel in complex reasoning tasks, and distilling their reasoning capabilities into smaller models has shown promise. However, we uncover an interesting phenomenon, which we term the Small Mod…