paper-with-me

홈 › Papers

Soft Knowledge Distillation with Multi-Dimensional Cross-Net Attention for Image Restoration Models Compression

2025-01-16 · Yongheng Zhang, Danfeng Yan

Transformer-based encoder-decoder models have achieved remarkable success in image-to-image transfer tasks, particularly in image restoration. However, their high computational complexity-manifested in elevated FLOPs and parameter counts-limits their application in real-world scenarios. Existing knowledge distillation methods in image restoration typically employ lightweight student models that directly mimic the intermediate features and reconstruction results of the teacher, overlooking the implicit attention relationships between them. To address this, we propose a Soft Knowledge Distillation (SKD) strategy that incorporates a Multi-dimensional Cross-net Attention (MCA) mechanism for compressing image restoration models. This mechanism facilitates interaction between the student and teacher across both channel and spatial dimensions, enabling the student to implicitly learn the attention matrices. Additionally, we employ a Gaussian kernel function to measure the distance between student and teacher features in kernel space, ensuring stable and efficient feature learning. To further enhance the quality of reconstructed images, we replace the commonly used L1 or KL divergence loss with a contrastive learning loss at the image level. Experiments on three tasks-image deraining, deblurring, and denoising-demonstrate that our SKD strategy significantly reduces computational complexity while maintaining strong image restoration capabilities.

📄 PDF Abstract BibTeX arXiv:2501.09321

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDeblurringDecoderDenoisingImage RestorationKnowledge DistillationRain Removal

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Contrastive Learning 설명 없음
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

DualDE: Dually Distilling Knowledge Graph Embedding for Faster and Cheaper Reasoning

2020-09-13 · Yushan Zhu, Wen Zhang, Mingyang Chen, Hui Chen 외

Knowledge Graph Embedding (KGE) is a popular method for KG reasoning and training KGEs with higher dimension are usually preferred since they have better reasoning capability. However, high-dimensional KGEs pose huge cha…

Graph EmbeddingKnowledge DistillationKnowledge Graph EmbeddingKnowledge Graph Embeddings+1

C2KD: Bridging the Modality Gap for Cross-Modal Knowledge Distillation

2024-01-01 · CVPR 2024 1 · Fushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang 외

Existing Knowledge Distillation (KD) methods typically focus on transferring knowledge from a large-capacity teacher to a low-capacity student model achieving substantial success in unimodal knowledge transfer. Howev…

Knowledge DistillationTransfer Learning

Balancing Knowledge Distillation for Imbalance Learning with Bilevel Optimization

2026-05-18 · Anh B. H. Nguyen, Ba Tho Phan, Viet Cuong Ta arxiv

Knowledge distillation transfers knowledge from a high capacity teacher to a compact student using a mixture of hard and soft losses. On imbalanced data, a fixed weighting between hard and soft losses becomes brittle the…

Knowledge DistillationBilevel Optimization

Aligning Logits Generatively for Principled Black-Box Knowledge Distillation

2022-05-21 · CVPR 2024 1 · Jing Ma, Xiang Xiang, Ke Wang, Yuchuan Wu 외

Black-Box Knowledge Distillation (B2KD) is a formulated problem for cloud-to-edge model compression with invisible data and models hosted on the server. B2KD faces challenges such as limited Internet exchange and edge-cl…

Federated LearningKnowledge DistillationModel Compression

Self-distillation with Batch Knowledge Ensembling Improves ImageNet Classification

2021-04-27 · Yixiao Ge, Xiao Zhang, Ching Lam Choi, Ka Chun Cheung 외

The recent studies of knowledge distillation have discovered that ensembling the "dark knowledge" from multiple teachers or students contributes to creating better soft targets for training, but at the cost of significan…

ClassificationGeneral ClassificationKnowledge Distillation