Papers Dataset Distillation
“Dataset Distillation” 태그가 달린 논문 216편 · 필터 해제
Information-Guided Diffusion Sampling for Dataset Distillation
Dataset distillation aims to create a compact dataset that retains essential information while maintaining model performance. Diffusion models (DMs) have shown promise for this task but struggle in low images-per-class (…
Dataset DistillationTask-Specific Generative Dataset Distillation with Difficulty-Guided Sampling
To alleviate the reliance of deep neural networks on large-scale datasets, dataset distillation aims to generate compact, high-quality synthetic datasets that can achieve comparable performance to the original dataset. T…
Dataset DistillationFADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
Residual connection has been extensively studied and widely applied at the model architecture level. However, its potential in the more challenging data-centric approaches remains unexplored. In this work, we introduce t…
Computational EfficiencyDataset DistillationGPUDataset Distillation via Vision-Language Category Prototype
Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consump…
Dataset DistillationDescriptiveLarge Language ModelCaO$_2$: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation
The recent introduction of diffusion models in dataset distillation has shown promising potential in creating compact surrogate datasets for large, high-resolution target datasets, offering improved efficiency and perfor…
Dataset DistillationFedWSIDD: Federated Whole Slide Image Classification via Dataset Distillation
Federated learning (FL) has emerged as a promising approach for collaborative medical image analysis, enabling multiple institutions to build robust predictive models while preserving sensitive patient data. In the conte…
ClassificationDataset DistillationFederated Learningimage-classification+2Dataset distillation for memorized data: Soft labels can leak held-out teacher knowledge
Dataset distillation aims to compress training data into fewer examples via a teacher, from which a student can learn effectively. While its success is often attributed to structure in the data, modern neural networks al…
Dataset DistillationMemorizationFlowing Datasets with Wasserstein over Wasserstein Gradient Flows
Many applications in machine learning involve data represented as probability distributions. The emergence of such data requires radically novel techniques to design tractable gradient flows on probability distributions …
Dataset DistillationDomain AdaptationTransfer LearningOD3: Optimization-free Dataset Distillation for Object Detection
Training large neural networks on large-scale datasets requires substantial computational resources, particularly for dense prediction tasks such as object detection. Although dataset distillation (DD) has been proposed …
Dataset Distillationimage-classificationImage Classificationobject-detection+1Hyperbolic Dataset Distillation
To address the computational and storage challenges posed by large-scale datasets in deep learning, dataset distillation has been proposed to synthesize a compact dataset that replaces the original while maintaining comp…
Computational EfficiencyDataset DistillationDiversity-Driven Generative Dataset Distillation Based on Diffusion Model with Self-Adaptive Memory
Dataset distillation enables the training of deep neural networks with comparable performance in significantly reduced time by compressing large datasets into small and representative ones. Although the introduction of g…
Dataset DistillationDiversityData-Distill-Net: A Data Distillation Approach Tailored for Reply-based Continual Learning
Replay-based continual learning (CL) methods assume that models trained on a small subset can also effectively minimize the empirical risk of the complete dataset. These methods maintain a memory buffer that stores a sam…
Continual LearningDataset DistillationMGD$^3$: Mode-Guided Dataset Distillation using Diffusion Models
Dataset distillation has emerged as an effective strategy, significantly reducing training costs and facilitating more efficient model deployment. Recent advances have leveraged generative models to distill datasets by c…
Dataset DistillationDiversityCONCORD: Concept-Informed Diffusion for Dataset Distillation
Dataset distillation (DD) has witnessed significant progress in creating small datasets that encapsulate rich information from large original ones. Particularly, methods based on generative priors show promising performa…
Computational EfficiencyDataset DistillationDenoisingImage GenerationTaming Diffusion for Dataset Distillation with High Representativeness
Recent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of d…
Dataset DistillationImage GenerationContrastive Learning-Enhanced Trajectory Matching for Small-Scale Dataset Distillation
Deploying machine learning models in resource-constrained environments, such as edge devices or rapid prototyping scenarios, increasingly demands distillation of large datasets into significantly smaller yet informative …
Contrastive LearningDataset DistillationImage GenerationExploring Generalized Gait Recognition: Reducing Redundancy and Noise within Indoor and Outdoor Datasets
Generalized gait recognition, which aims to achieve robust performance across diverse domains, remains a challenging problem due to severe domain shifts in viewpoints, appearances, and environments. While mixed-dataset t…
Dataset DistillationGait RecognitionRepresentation LearningTripletDD-Ranking: Rethinking the Evaluation of Dataset Distillation
In recent years, dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance comparable to those trained on the origina…
Data AugmentationData CompressionDataset DistillationBeyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation
Multimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approach…
cross-modal alignmentDataset DistillationLeveraging Multi-Modal Information to Enhance Dataset Distillation
Dataset distillation aims to create a compact and highly representative synthetic dataset that preserves the knowledge of a larger real dataset. While existing methods primarily focus on optimizing visual representations…
Dataset DistillationObject