Papers Self-Knowledge Distillation
“Self-Knowledge Distillation” 태그가 달린 논문 68편 · 필터 해제
Tackling Data Heterogeneity in Federated Learning through Knowledge Distillation with Inequitable Aggregation
Federated learning aims to train a global model in a distributed environment that is close to the performance of centralized training. However, issues such as client label skew, data quantity skew, and other heterogeneit…
Federated LearningKnowledge DistillationSelf-Knowledge DistillationMoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
Multimodal Large Language Models (MLLMs) excel in understanding complex language and visual data, enabling generalist robotic systems to interpret instructions and perform embodied tasks. Nevertheless, their real-world d…
Knowledge DistillationMixture-of-ExpertsRobot ManipulationSelf-Knowledge Distillation+1xVLM2Vec: Adapting LVLM-based embedding models to multilinguality using Self-Knowledge Distillation
In the current literature, most embedding models are based on the encoder-only transformer architecture to extract a dense and meaningful representation of the given input, which can be a text, an image, and more. With t…
Knowledge DistillationLanguage ModelingLanguage ModellingSelf-Knowledge DistillationInvestigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models
Omnimodal Large Language Models (OLLMs) have shown significant progress in integrating vision and text, but still struggle with integrating vision and audio, often exhibiting suboptimal performance when processing audio …
Knowledge DistillationSelf-Knowledge DistillationEfficient Lung Ultrasound Severity Scoring Using Dedicated Feature Extractor
With the advent of the COVID-19 pandemic, ultrasound imaging has emerged as a promising technique for COVID-19 detection, due to its non-invasive nature, affordability, and portability. In response, researchers have focu…
DiagnosticKnowledge DistillationSelf-Knowledge DistillationVideo ClassificationGenerative Dataset Distillation Based on Self-knowledge Distillation
Dataset distillation is an effective technique for reducing the cost and complexity of model training while maintaining performance by compressing large datasets into smaller, more efficient versions. In this paper, we p…
Dataset DistillationKnowledge DistillationSelf-Knowledge DistillationTowards Satellite Non-IID Imagery: A Spectral Clustering-Assisted Federated Learning Approach
Low Earth orbit (LEO) satellites are capable of gathering abundant Earth observation data (EOD) to enable different Internet of Things (IoT) applications. However, to accomplish an effective EOD processing mechanism, it …
Earth ObservationFederated LearningKnowledge DistillationSelf-Knowledge DistillationFrequency-Guided Masking for Enhanced Vision Self-Supervised Learning
We present a novel frequency-based Self-Supervised Learning (SSL) approach that significantly enhances its efficacy for pre-training. Prior work in this direction masks out pre-defined frequencies in the input image and …
Few-Shot Learningimage-classificationImage ClassificationImage Compression+4SalNAS: Efficient Saliency-prediction Neural Architecture Search with self-knowledge distillation
Recent advancements in deep convolutional neural networks have significantly improved the performance of saliency prediction. However, the manual configuration of the neural network architectures requires domain knowledg…
DecoderKnowledge DistillationNeural Architecture SearchPrediction+2Towards A Generalizable Pathology Foundation Model via Unified Knowledge Distillation
Foundation models pretrained on large-scale datasets are revolutionizing the field of computational pathology (CPath). The generalization ability of foundation models is crucial for the success in various downstream clin…
Knowledge DistillationQuestion AnsweringRepresentation LearningSelf-Knowledge Distillation+3Three-Stream Temporal-Shift Attention Network Based on Self-Knowledge Distillation for Micro-Expression Recognition
Micro-expressions are subtle facial movements that occur spontaneously when people try to conceal real emotions. Micro-expression recognition is crucial in many fields, including criminal analysis and psychotherapy. Howe…
Knowledge DistillationMicro Expression RecognitionMicro-Expression RecognitionMotion Magnification+1SeCoKD: Aligning Large Language Models for In-Context Learning with Fewer Shots
Previous studies have shown that demonstrations can significantly help Large Language Models (LLMs ) perform better on the given tasks. However, this so-called In-Context Learning ( ICL ) ability is very sensitive to the…
In-Context LearningKnowledge DistillationSelf-Knowledge DistillationSelf-Knowledge Distillation for Learning Ambiguity
Recent language models have shown remarkable performance on natural language understanding (NLU) tasks. However, they are often sub-optimal when faced with ambiguous samples that can be interpreted in multiple ways, over…
Knowledge DistillationNatural Language UnderstandingSelf-Knowledge DistillationGuiding Frame-Level CTC Alignments Using Self-knowledge Distillation
Transformer encoder with connectionist temporal classification (CTC) framework is widely used for automatic speech recognition (ASR). However, knowledge distillation (KD) for ASR displays a problem of disagreement betwee…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge DistillationSelf-Knowledge Distillation+2Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
Skeleton-based action representation learning aims to interpret and understand human behaviors by encoding the skeleton sequences, which can be categorized into two primary training paradigms: supervised learning and sel…
Action RecognitionContrastive LearningKnowledge DistillationRepresentation Learning+4CrossMatch: Enhance Semi-Supervised Medical Image Segmentation with Perturbation Strategies and Knowledge Distillation
Semi-supervised learning for medical image segmentation presents a unique challenge of efficiently using limited labeled data while leveraging abundant unlabeled data. Despite advancements, existing methods often do not …
Image SegmentationKnowledge DistillationMedical Image SegmentationSelf-Knowledge Distillation+2Weakly Supervised Monocular 3D Detection with a Single-View Image
Monocular 3D detection (M3D) aims for precise 3D object localization from a single-view image which usually involves labor-intensive annotation of 3D detection boxes. Weakly supervised M3D has recently been studied to ob…
Knowledge DistillationObject LocalizationSelf-Knowledge DistillationTransfer LearningDistilled Gradual Pruning with Pruned Fine-tuning
Neural Networks (NNs) have been driving machine learning progress in recent years, but their larger models present challenges in resource-limited environments. Weight pruning reduces the computational demand, often with …
Image ClassificationKnowledge DistillationSelf-Knowledge DistillationBGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
In this paper, we present a new embedding model, called M3-Embedding, which is distinguished for its versatility in Multi-Linguality, Multi-Functionality, and Multi-Granularity. It can support more than 100 working langu…
Knowledge DistillationRetrievalSelf-Knowledge DistillationDeep Clustering with Diffused Sampling and Hardness-aware Self-distillation
Deep clustering has gained significant attention due to its capability in learning clustering-friendly representations without labeled data. However, previous deep clustering methods tend to treat all samples equally, wh…
ClusteringContrastive LearningDeep ClusteringKnowledge Distillation+2