paper-with-me

홈 › Papers

Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer

2026-05-12 · Nilushika Udayangani, Kishor Nandakishor, Marimuthu Palaniswami arxiv

While traditional time-series classifiers assume full sequences at inference, practical constraints (latency and cost) often limit inputs to partial prefixes. The absence of class-discriminative patterns in partial data can significantly hinder a classifier's ability to generalize. This work uses knowledge distillation (KD) to equip partial time series classifiers with the generalization ability of their full-sequence counterparts. In KD, high-capacity teacher transfers supervision to aid student learning on the target task. Matching with teacher features has shown promise in closing the generalization gap due to limited parameter capacity. However, when the generalization gap arises from training-data differences (full versus partial), the teacher's full-context features can be an overwhelming target signal for the student's short-context features. To provide progressive, diverse, and collective teacher supervision, we propose Generative Diffusion Prior Distillation (GDPD), a novel KD framework that treats short-context student features as degraded observations of the target full-context features. Inspired by the iterative restoration capability of diffusion models, we learn a diffusion-based generative prior over teacher features. Leveraging this prior, we posterior-sample target teacher representations that could best explain the missing long-range information in the student features and optimize the student features to be minimally degraded relative to these targets. GDPD provides each student feature with a distribution of task-relevant long-context knowledge, which benefits learning on the partial classification task. Extensive experiments across earliness settings, datasets, and architectures demonstrate GDPD's effectiveness for full-to-partial distillation.

📄 PDF Abstract BibTeX arXiv:2605.11414

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

SER-Diff: Synthetic Error Replay Diffusion for Incremental Brain Tumor Segmentation

2025-10-06 · Sashank Makanaboyina arxiv

Incremental brain tumor segmentation is critical for models that must adapt to evolving clinical datasets without retraining on all prior data. However, catastrophic forgetting, where models lose previously acquired know…

Brain Tumor SegmentationKnowledge DistillationIncremental Learning

One-Step Diffusion-Based Image Compression with Semantic Distillation

2025-05-22 · Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li 외

While recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasing latency. In this work, we revisit the design of a diffusion-based codec and a…

Image Compression

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

2024-11-27 · CVPR 2025 1 · Shengqu Cai, Eric Chan, Yunzhi Zhang, Leonidas Guibas 외

Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images of a specific instance in novel contexts…

Image GenerationLanguage ModelingLanguage Modelling

Pool-Select-Refine for Allocation-Aware Generative Dataset Distillation

2026-06-01 · Wenmin Li, Shunsuke Sakai, Zhongkai Zhao, Tatsuhito Hasegawa arxiv

Diffusion-based dataset distillation has recently emerged as a promising paradigm for condensing large-scale datasets into compact synthetic sets. By leveraging pretrained generative priors, these methods can produce rea…

Fine-Grained Image Classification

Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond

2025-05-07 · Jessie Richter-Powell, Antonio Torralba, Jonathan Lorraine

We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of…

3D GenerationAudio Source SeparationText to 3D