paper-with-me

홈 › Papers

Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation

2026-03-31 · Hoang-Chau Luong, Dat Ba Tran, Lingwei Chen arxiv

Reverse Kullback-Leibler (RKL) divergence has recently emerged as the preferred objective for large language model (LLM) distillation, consistently outperforming forward KL (FKL), particularly in regimes with large vocabularies and significant teacher-student capacity mismatch, where RKL focuses learning on dominant modes rather than enforcing dense alignment. However, RKL introduces a structural limitation that drives the student toward overconfident predictions. We first provide an analysis of RKL by decomposing its gradients into target and non-target components, and show that non-target gradients consistently push the target logit upward even when the student already matches the teacher, thereby reducing output diversity. In addition, RKL provides weak supervision over non-target classes, leading to poor tail alignment. To address these issues, we propose Diversity-aware RKL (DRKL), which removes this gradient effect and strengthens non-target supervision while preserving the optimization benefits of RKL. Extensive experiments across datasets and model families demonstrate that DRKL consistently outperforms FKL, RKL, and other state-of-the-art distillation objectives, achieving better performance and a superior fidelity-diversity trade-off.

📄 PDF Abstract BibTeX arXiv:2604.00223

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models

2024-04-03 · Taiqiang Wu, Chaofan Tao, Jiahao Wang, Runming Yang 외

Kullback-Leiber divergence has been widely used in Knowledge Distillation (KD) to compress Large Language Models (LLMs). Contrary to prior assertions that reverse Kullback-Leibler (RKL) divergence is mode-seeking and thu…

DiversityKnowledge Distillation

Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization

2024-09-15 · Haoyuan Sun, Bo Xia, Yongzhe Chang, Xueqian Wang

Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable i…

DiversityImage GenerationText to Image GenerationText-to-Image Generation

Improved Training of Generative Adversarial Networks Using Representative Features

2018-01-28 · ICML 2018 7 · Duhyeon Bang, Hyunjung Shim

Despite the success of generative adversarial networks (GANs) for image generation, the trade-off between visual quality and image diversity remains a significant issue. This paper achieves both aims simultaneously by im…

DiversityImage Generation

Ratio Divergence Learning Using Target Energy in Restricted Boltzmann Machines: Beyond Kullback--Leibler Divergence Learning

2024-09-12 · Yuichi Ishida, Yuma Ichikawa, Aki Dote, Toshiyuki Miyazawa 외

We propose ratio divergence (RD) learning for discrete energy-based models, a method that utilizes both training data and a tractable target energy function. We apply RD learning to restricted Boltzmann machines (RBMs), …

On Voronoi diagrams and dual Delaunay complexes on the information-geometric Cauchy manifolds

2020-06-12 · Frank Nielsen

We study the Voronoi diagrams of a finite set of Cauchy distributions and their dual complexes from the viewpoint of information geometry by considering the Fisher-Rao distance, the Kullback-Leibler divergence, the chi s…