Knowledge Distillation
7개 벤치마크 · 논문 5,285편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Focal Loss for Dense Object Detection
Distilling the Knowledge in a Neural Network
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
Papers
On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data
Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from …
Knowledge DistillationDeep Neural Networks for Learning Intent from sEMG Signals to Support Hardware Devices for Post-Stroke Neurorehabilitation
Finger-specific motor intent is a clinically meaningful control signal for post-stroke neurorehabilitation, where residual muscle activity may remain measurable despite weak or incomplete movement. We study five-finger m…
Knowledge DistillationPretraining and Distillation Matter More Than Architecture Family for Label-Free Single-Cell Classification
Choosing a deep learning architecture for label-free single-cell classification remains an open question, with microscopy benchmarks reporting conflicting conclusions about CNNs versus transformers. We present a controll…
Knowledge DistillationPersistent Teacher Anchoring for Tool-Using Agents
Distillation is common in LLM post-training, where on-policy knowledge distillation (OPKD) uses student-generated trajectories to prepare the student for downstream RL. At each state, the student matches a next-token dis…
Knowledge DistillationImportance-Aware Low-Rank Distillation of Diffusion Transformers
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generation, yet their scale poses challenges for efficient deployment. While truncated singular value decomposition (SV…
Text-to-Image GenerationKnowledge DistillationKnowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through contr…
Self-Supervised LearningKnowledge Distillation