paper-with-me

홈 › Papers

Information-Preserving Reformulation of Reasoning Traces for Antidistillation

2025-10-13 · Jiayu Ding, Lei Cui, Li Dong, Nanning Zheng, Furu Wei arxiv

Recent advances in Large Language Models (LLMs) show that extending the length of reasoning chains significantly improves performance on complex tasks. While revealing these reasoning traces helps users better follow, verify, and learn from the model's problem-solving process, it also makes them highly vulnerable to unauthorized distillation. To mitigate this risk, proprietary model providers often adopt aggressive protection strategies, such as replacing detailed reasoning with brief summaries, which deprive users of valuable intermediate information. To address this trade-off, we propose PART, an information-preserving antidistillation reformulation of reasoning traces. Motivated by the difference between how humans understand reasoning traces and how LLMs exploit them for supervised fine-tuning, we design a simple but effective two-step reformulation: removing self-talk behaviors and reordering sub-conclusions. A small auxiliary model is trained to perform this reformulation, incurring minimal computational overhead. Extensive experiments demonstrate that PART consistently disrupts distillation across student models of different sizes and types on various reasoning benchmarks. For instance, when training on reformulated traces, even the performance of a large 32B student model decreases from 54.17 to 46.88 on AIME 2024, corresponding to a 13.5% degradation.

📄 PDF Abstract BibTeX arXiv:2510.11545

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Antidistillation Sampling

2025-04-17 · Yash Savani, Asher Trockman, Zhili Feng, Avi Schwarzschild 외

Frontier models that generate extended reasoning traces inadvertently produce rich token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that li…

Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models

2026-04-25 · Max Hartman, Vidhata Jayaraman, Moulik Choraria, Yash Savani 외 arxiv

Distillation via sampling reasoning traces exposes closed-source frontier models to adversarial third parties who can bypass their guardrails and misappropriate their capabilities. Antidistillation methods aim to address…

Asking Back: Interaction-Layer Antidistillation Watermarks

2026-05-15 · Guang Yang, Amir Ghasemian, Fengchen Liu, Zhong Wang 외 arxiv

Detecting unauthorized knowledge distillation from a deployed LLM API is hard because the defender controls neither the attacker's training pipeline nor the next-token logits. Existing defenses operate on the teacher's o…

Knowledge Distillation

Antidistillation Fingerprinting

2026-02-03 · Yixuan Even Xu, John Kirchenbauer, Yash Savani, Asher Trockman 외 arxiv

Model distillation enables efficient emulation of frontier large language models (LLMs), creating a need for robust mechanisms to detect when a third-party student model has trained on a teacher model's outputs. However,…

Mathematical ReasoningCode Generation

The Distillation Game: Adaptive Attacks & Efficient Defenses

2026-05-21 · Youssef Allouah, Mahdi Haghifam, Sanmi Koyejo, Reza Shokri arxiv

Distillation attacks create a deployment trade-off for model providers: the same outputs that make a model more useful can also make it easier to imitate. We study this trade-off through a minimax game between a utility-…