paper-with-me

홈 › Papers

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models

2026-05-28 · Qi Liu, Mingdi Sun, Yongyi He, Zhi Zheng, Tong Xu, Yi Zheng, Zhefeng Wang, Enhong Chen arxiv

Supervised fine-tuning (SFT) followed by reinforcement learning (RL) has become a standard post-training paradigm for large language models. This paradigm provides a cold-start for RL exploration, avoiding the inefficiency of pure RL where on-policy sampling yields insufficient positive samples. However, in practice, existing approaches often use a small amount of data for SFT initialization compared to the RL phase, which can cause the model to fit the limited samples and shift away from its pre-trained distribution. This distribution shift impedes the model's ability to effectively explore during subsequent RL training. To address this challenge, we propose that in low-data regimes, SFT should prioritize activating task-relevant capabilities rather than memorizing specific content. Along this line, we propose EKSFT (Entropy-KL Selective Fine-Tuning), which selectively masks tokens that exhibit either high entropy or high KL divergence from a reference model. By excluding these high-uncertainty, distribution-shifting tokens from imitation, EKSFT injects task-specific knowledge while preserving the integrity of the model's pre-trained distribution. Empirical evaluations on mathematical reasoning benchmarks demonstrate that EKSFT consistently outperforms standard SFT. Further RL fine-tuning from the EKSFT model yields consistently better post-RL performance, indicating improved exploration for the RL stage. Our codes and datasets are available at https://github.com/MINE-USTC/EKSFT.

📄 PDF Abstract BibTeX arXiv:2605.29303

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

2026-07-12 · Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan 외 arxiv

On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In th…

Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR

2023-11-08 · Qian Chen, Wen Wang, Qinglin Zhang, Siqi Zheng 외

Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and …

Decoder

Train No Evil: Selective Masking for Task-Guided Pre-Training

2020-04-21 · EMNLP 2020 11 · Yuxian Gu, Zhengyan Zhang, Xiaozhi Wang, Zhiyuan Liu 외

Recently, pre-trained language models mostly follow the pre-train-then-fine-tuning paradigm and have achieved great performance on various downstream tasks. However, since the pre-training stage is typically task-agnosti…

Language ModelingLanguage ModellingMasked Language ModelingSentiment Analysis

EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models

2026-03-19 · Minsoo Cheong, Donghyun Son, Woosang Lim, Sungjoo Yoo arxiv

Diffusion-based large language models (dLLMs) rely on bidirectional attention, which prevents lossless KV caching and requires a full forward pass at every denoising step. Existing approximate KV caching methods reduce t…

VISTA: Auditing Semantic Divergence in Vision-Language Models

2026-07-03 · Junchi Liao, Jiawen Deng, Fuji Ren arxiv

Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce unusually uniform responses that differ…