paper-with-me

홈 › Papers

ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks

2026-05-31 · Dhruv Saini, Rohan Pandey arxiv

Large language models often improve on difficult tasks by spending inference-time compute on a reasoning trace before producing the final answer. That extra computation can be useful, but it also raises latency, token cost, and deployment complexity. We introduce \textbf{ThinkSwitch}, a low-compute procedure for co-training paired instruct and thinking checkpoints. Starting from compatible Qwen3-4B instruct and thinking models, each iteration asks the thinking checkpoint to generate answers, removes the reasoning trace, distills the answer-only pairs into the instruct checkpoint with QLoRA, and reconstructs a thinking checkpoint with spherical weight interpolation. The only human-supplied inputs are task prompts; the labels are generated by the model itself. On a 30-question AIME 2026 evaluation, ThinkSwitch improves the instruct checkpoint from 10/30 to 20/30 and the thinking checkpoint from 14/30 to 22/30. On a 30-question PubMedQA subset, it improves the instruct checkpoint from 13/30 to 18/30 and the thinking checkpoint from 18/30 to 25/30. The complete experiment uses 15 training prompts per domain and costs \$2.86 on a single cloud RTX 3070. The results are small-scale, but they indicate that targeted distillation loops can move part of the benefit of explicit reasoning into weights while preserving a separate thinking mode.

📄 PDF Abstract BibTeX arXiv:2606.01080

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ThinkSwitcher: When to Think Hard, When to Think Fast

2025-05-20 · Guosheng Liang, Longguang Zhong, ZiYi Yang, Xiaojun Quan

Large reasoning models (LRMs) excel at solving complex tasks by leveraging long chain-of-thought (CoT) reasoning. However, this often leads to overthinking on simple tasks, resulting in unnecessary computational overhead…

LongQLoRA: Efficient and Effective Method to Extend Context Length of Large Language Models

2023-11-08 · Jianxin Yang

We present LongQLoRA, an efficient and effective method to extend context length of large language models with less training resources. LongQLoRA combines the advantages of Position Interpolation, QLoRA and Shift Short A…

8kGPU

PAINT: Partial-Solution Adaptive Interpolated Training for Self-Distilled Reasoners

2026-04-29 · Zhiquan Tan, Yinrong Hong arxiv

Improving large language model (LLM) reasoning requires supervision that is both aligned with the model's own test-time states and informative at the token level. Reinforcement learning with verifiable rewards provides o…

Reinforcement Learning

Doc-to-Atom: Learning to Compile and Compose Memory Atoms

2026-06-10 · Xingjian Diao, Wenbo Li, Yashas Malur Saidutta, Avinash Amballa 외 arxiv

Long input sequences are central to document understanding and multi-step reasoning in Large Language Models, yet the quadratic cost of attention makes inference both memory-intensive and slow. Context distillation mitig…

Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context

2026-05-31 · Shihao Ji, Mingyu Li, Zihui Song arxiv

The quadratic complexity of self-attention remains a bottleneck for Large Language Models (LLMs) processing ultra-long contexts. The Naive Bayes Cognitive Engine (NBCE) parallelizes long-context inference by chunking doc…