paper-with-me

Papers Dataset Distillation

“Dataset Distillation” 태그가 달린 논문 216편 · 필터 해제

Information-Guided Diffusion Sampling for Dataset Distillation

2025-07-07 · Linfeng Ye, Shayan Mohajer Hamidi, Guang Li, Takahiro Ogawa 외

Dataset distillation aims to create a compact dataset that retains essential information while maintaining model performance. Diffusion models (DMs) have shown promise for this task but struggle in low images-per-class (…

Dataset Distillation

Task-Specific Generative Dataset Distillation with Difficulty-Guided Sampling

2025-07-04 · Mingzhuo Li, Guang Li, Jiafeng Mao, Linfeng Ye 외

To alleviate the reliance of deep neural networks on large-scale datasets, dataset distillation aims to generate compact, high-quality synthetic datasets that can achieve comparable performance to the original dataset. T…

Dataset Distillation

FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation

2025-06-30 · Jiacheng Cui, Xinyue Bi, Yaxin Luo, Xiaohan Zhao 외

Residual connection has been extensively studied and widely applied at the model architecture level. However, its potential in the more challenging data-centric approaches remains unexplored. In this work, we introduce t…

Computational EfficiencyDataset DistillationGPU

Dataset Distillation via Vision-Language Category Prototype

2025-06-30 · Yawen Zou, Guang Li, Duo Su, Zi Wang 외

Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consump…

Dataset DistillationDescriptiveLarge Language Model

CaO$_2$: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation

2025-06-27 · Haoxuan Wang, Zhenghao Zhao, Junyi Wu, Yuzhang Shang 외

The recent introduction of diffusion models in dataset distillation has shown promising potential in creating compact surrogate datasets for large, high-resolution target datasets, offering improved efficiency and perfor…

Dataset Distillation

FedWSIDD: Federated Whole Slide Image Classification via Dataset Distillation

2025-06-18 · Haolong Jin, Shenglin Liu, Cong Cong, Qingmin Feng 외

Federated learning (FL) has emerged as a promising approach for collaborative medical image analysis, enabling multiple institutions to build robust predictive models while preserving sensitive patient data. In the conte…

ClassificationDataset DistillationFederated Learningimage-classification+2

Dataset distillation for memorized data: Soft labels can leak held-out teacher knowledge

2025-06-17 · Freya Behrens, Lenka Zdeborová

Dataset distillation aims to compress training data into fewer examples via a teacher, from which a student can learn effectively. While its success is often attributed to structure in the data, modern neural networks al…

Dataset DistillationMemorization

Flowing Datasets with Wasserstein over Wasserstein Gradient Flows

2025-06-09 · Clément Bonet, Christophe Vauthier, Anna Korba

Many applications in machine learning involve data represented as probability distributions. The emergence of such data requires radically novel techniques to design tractable gradient flows on probability distributions …

Dataset DistillationDomain AdaptationTransfer Learning

OD3: Optimization-free Dataset Distillation for Object Detection

2025-06-02 · Salwa K. Al Khatib, Ahmed Elhagry, Shitong Shao, Zhiqiang Shen

Training large neural networks on large-scale datasets requires substantial computational resources, particularly for dense prediction tasks such as object detection. Although dataset distillation (DD) has been proposed …

Dataset Distillationimage-classificationImage Classificationobject-detection+1

Hyperbolic Dataset Distillation

2025-05-30 · Wenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa 외

To address the computational and storage challenges posed by large-scale datasets in deep learning, dataset distillation has been proposed to synthesize a compact dataset that replaces the original while maintaining comp…

Computational EfficiencyDataset Distillation

Diversity-Driven Generative Dataset Distillation Based on Diffusion Model with Self-Adaptive Memory

2025-05-26 · Mingzhuo Li, Guang Li, Jiafeng Mao, Takahiro Ogawa 외

Dataset distillation enables the training of deep neural networks with comparable performance in significantly reduced time by compressing large datasets into small and representative ones. Although the introduction of g…

Dataset DistillationDiversity

Data-Distill-Net: A Data Distillation Approach Tailored for Reply-based Continual Learning

2025-05-26 · Wenyang Liao, Quanziang Wang, Yichen Wu, Renzhen Wang 외

Replay-based continual learning (CL) methods assume that models trained on a small subset can also effectively minimize the empirical risk of the complete dataset. These methods maintain a memory buffer that stores a sam…

Continual LearningDataset Distillation

MGD$^3$: Mode-Guided Dataset Distillation using Diffusion Models

2025-05-25 · Jeffrey A. Chan-Santiago, Praveen Tirupattur, Gaurav Kumar Nayak, Gaowen Liu 외

Dataset distillation has emerged as an effective strategy, significantly reducing training costs and facilitating more efficient model deployment. Recent advances have leveraged generative models to distill datasets by c…

Dataset DistillationDiversity

CONCORD: Concept-Informed Diffusion for Dataset Distillation

2025-05-23 · Jianyang Gu, Haonan Wang, Ruoxi Jia, Saeed Vahidian 외

Dataset distillation (DD) has witnessed significant progress in creating small datasets that encapsulate rich information from large original ones. Particularly, methods based on generative priors show promising performa…

Computational EfficiencyDataset DistillationDenoisingImage Generation

Taming Diffusion for Dataset Distillation with High Representativeness

2025-05-23 · Lin Zhao, Yushu Wu, Xinru Jiang, Jianyang Gu 외

Recent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of d…

Dataset DistillationImage Generation

Contrastive Learning-Enhanced Trajectory Matching for Small-Scale Dataset Distillation

2025-05-21 · Wenmin Li, Shunsuke Sakai, Tatsuhito Hasegawa

Deploying machine learning models in resource-constrained environments, such as edge devices or rapid prototyping scenarios, increasingly demands distillation of large datasets into significantly smaller yet informative …

Contrastive LearningDataset DistillationImage Generation

Exploring Generalized Gait Recognition: Reducing Redundancy and Noise within Indoor and Outdoor Datasets

2025-05-21 · Qian Zhou, Xianda Guo, Jilong Wang, Chuanfu Shen 외

Generalized gait recognition, which aims to achieve robust performance across diverse domains, remains a challenging problem due to severe domain shifts in viewpoints, appearances, and environments. While mixed-dataset t…

Dataset DistillationGait RecognitionRepresentation LearningTriplet

DD-Ranking: Rethinking the Evaluation of Dataset Distillation

2025-05-19 · Zekai Li, Xinhao Zhong, Samir Khaki, Zhiyuan Liang 외

In recent years, dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance comparable to those trained on the origina…

Data AugmentationData CompressionDataset Distillation

Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation

2025-05-16 · Xin Zhang, Ziruo Zhang, Jiawei Du, Zuozhu Liu 외

Multimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approach…

cross-modal alignmentDataset Distillation

Leveraging Multi-Modal Information to Enhance Dataset Distillation

2025-05-13 · Zhe Li, Hadrien Reynaud, Bernhard Kainz

Dataset distillation aims to create a compact and highly representative synthetic dataset that preserves the knowledge of a larger real dataset. While existing methods primarily focus on optimizing visual representations…

Dataset DistillationObject
1–20 / 216 다음 →