paper-with-me

Knowledge Distillation

7개 벤치마크 · 논문 5,285편 · 이 태스크의 논문 보기 →

Benchmarks

ImageNet

결과 59개

CIFAR-100

결과 28개

COCO 2017 val

결과 3개

PASCAL VOC

결과 2개

Cityscapes

결과 1개

KITTI

결과 1개

Most implemented

Focal Loss for Dense Object Detection

2017-08-07 · 구현 234개

Papers

On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

2026-09-09 · Hongyuan Zhang, Xianda Guo, Yanlun Peng, Qianlong Yang 외 arxiv

Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from …

Knowledge Distillation

Deep Neural Networks for Learning Intent from sEMG Signals to Support Hardware Devices for Post-Stroke Neurorehabilitation

2026-09-09 · Zakariyya Brewster, Divy Wadhwani, Emily Yan, Aidan Wang 외 arxiv

Finger-specific motor intent is a clinically meaningful control signal for post-stroke neurorehabilitation, where residual muscle activity may remain measurable despite weak or incomplete movement. We study five-finger m…

Knowledge Distillation

Pretraining and Distillation Matter More Than Architecture Family for Label-Free Single-Cell Classification

2026-09-09 · Philip Graemer, Giuseppe Di Caprio arxiv

Choosing a deep learning architecture for label-free single-cell classification remains an open question, with microscopy benchmarks reporting conflicting conclusions about CNNs versus transformers. We present a controll…

Knowledge Distillation

Persistent Teacher Anchoring for Tool-Using Agents

2026-09-04 · Hyun Bin Park, Kyungho Song, Sangmin Lee, Du-Seong Chang arxiv

Distillation is common in LLM post-training, where on-policy knowledge distillation (OPKD) uses student-generated trajectories to prepare the student for downstream RL. At each state, the student matches a next-token dis…

Knowledge Distillation

Importance-Aware Low-Rank Distillation of Diffusion Transformers

2026-09-04 · Denis Zavadski, Sebastian Heid, Damjan Kalšan, Stefan Roth 외 arxiv

Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generation, yet their scale poses challenges for efficient deployment. While truncated singular value decomposition (SV…

Text-to-Image GenerationKnowledge Distillation

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

2026-09-01 · Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li 외 hf

Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through contr…

Self-Supervised LearningKnowledge Distillation

전체 5,285편 보기 →