paper-with-me

홈 › Papers

RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting

2026-05-29 · Yuduo Li, Xiaofeng Shi, Qian Kou, Longbin Yu, Hua Zhou arxiv

Domain-specific supervised fine-tuning (SFT) often improves in-domain performance at the cost of degrading a model's general capabilities. We view this degradation through two practical gaps in domain SFT: a supervision-compatibility gap, where domain targets differ in style and reasoning format from the original model's natural responses, and a trajectory-preservation gap, where teacher-forced SFT optimizes fixed target tokens without constraining the model's behavior on its own generated prefixes. This process fails to preserve the model's original behavior. We propose RAFT (Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting), a two-stage framework that addresses both factors. First, RAFT constructs model-compatible supervision through self-conditioned rewriting, semantic filtering, and answer fusion. Second, RAFT performs Answer-Conditioned On-Policy Distillation, where the original instruction-tuned model provides soft targets on student-generated trajectories while being conditioned on the fused answer as helpful context. We further introduce top-K temperature distillation and EMA-based adaptive loss balancing to stabilize the domain-general trade-off. Across three instruction-tuned backbones and five domains, RAFT improves average domain accuracy by 23.2% over standard SFT, while recovering part of the SFT-induced degradation on MS-Bench and IFEval, with relative improvements of 18.2% and 10.2%, respectively. These results show that coupling data refinement with trajectory-level preservation provides an effective recipe for domain fine-tuning with alleviated forgetting.

📄 PDF Abstract BibTeX arXiv:2606.00147

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training Domain Draft Models for Speculative Decoding: Best Practices and Insights

2025-03-10 · Fenglu Hong, Ravi Raju, Jonathan Lingjie Li, Bo Li 외

Speculative decoding is an effective method for accelerating inference of large language models (LLMs) by employing a small draft model to predict the output of a target model. However, when adapting speculative decoding…

Knowledge Distillation

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation

2026-05-30 · Zining Liu, Yunhai Hu, Tianhua Xia, Bo Bao 외 arxiv

Speculative decoding (SD) has proven to be an effective technique for accelerating autoregressive generation in large language models (LLMs) however, its application to vision-language models (VLMs) remains relatively un…

Neural Architecture Searchmultimodal generation

Distillation-Guided Image Inpainting

2021-01-01 · ICCV 2021 10 · Maitreya Suin, Kuldeep Purohit, A. N. Rajagopalan

Image inpainting methods have shown significant improvements by using deep neural networks recently. However, many of these techniques often create distorted structures or blurry inconsistent textures. The problem is…

Image Inpainting

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

2026-08-14 · Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang 외 arxiv

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interacti…

Domain AdaptationVideo Generation

Forward-Free Diffusion Language Models with BPTT-Free Looped Refinement

2026-06-06 · Haotian Sun, Rushi Qiang, Yuqian Zheng, Bo Dai arxiv

Diffusion language models generate text through iterative denoising, offering a powerful alternative to autoregressive generation. However, discrete language spaces lack a natural neighborhood structure for defining effe…