paper-with-me

홈 › Papers

DomainStudio: Fine-Tuning Diffusion Models for Domain-Driven Image Generation using Limited Data

2023-06-25 · Jingyuan Zhu, Huimin Ma, Jiansheng Chen, Jian Yuan

Denoising diffusion probabilistic models (DDPMs) have been proven capable of synthesizing high-quality images with remarkable diversity when trained on large amounts of data. Typical diffusion models and modern large-scale conditional generative models like text-to-image generative models are vulnerable to overfitting when fine-tuned on extremely limited data. Existing works have explored subject-driven generation using a reference set containing a few images. However, few prior works explore DDPM-based domain-driven generation, which aims to learn the common features of target domains while maintaining diversity. This paper proposes a novel DomainStudio approach to adapt DDPMs pre-trained on large-scale source datasets to target domains using limited data. It is designed to keep the diversity of subjects provided by source domains and get high-quality and diverse adapted samples in target domains. We propose to keep the relative distances between adapted samples to achieve considerable generation diversity. In addition, we further enhance the learning of high-frequency details for better generation quality. Our approach is compatible with both unconditional and conditional diffusion models. This work makes the first attempt to realize unconditional few-shot image generation with diffusion models, achieving better quality and greater diversity than current state-of-the-art GAN-based approaches. Moreover, this work also significantly relieves overfitting for conditional generation and realizes high-quality domain-driven generation, further expanding the applicable scenarios of modern large-scale text-to-image models.

📄 PDF Abstract BibTeX arXiv:2306.14153

Code (1)

bbzhu-jy16/DomainStudio pytorch

Tasks

DenoisingDiversityImage Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

TF-ICON: Diffusion-Based Training-Free Cross-Domain Image Composition

2023-07-24 · ICCV 2023 1 · Shilin Lu, Yanzhu Liu, Adams Wai-Kin Kong

Text-driven diffusion models have exhibited impressive generative capabilities, enabling various image editing tasks. In this paper, we propose TF-ICON, a novel Training-Free Image COmpositioN framework that harnesses th…

Image-Guided CompositionText-to-Image Generation

Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment

2024-06-06 · CVPR 2025 1 · Jiayi Guo, Junhao Zhao, Chaoqun Du, Yulin Wang 외

Test-time adaptation (TTA) aims to improve the performance of source-domain pre-trained models on previously unseen, shifted target domains. Traditional TTA methods primarily adapt model weights based on target data stre…

Test-time Adaptation

Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning

2026-07-23 · Yi Xiong, Yuan-Yuan Cheng, Xiao-Ming Fu arxiv

Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained model's broad generative capability. Existing full and parameter-effic…

parameter-efficient fine-tuning

DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric Finetuning

2024-11-07 · Yuxuan Duan, Yan Hong, Bo Zhang, Jun Lan 외

The recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of thes…

AttributeDisentanglementImage Generation

Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation

2025-03-13 · Yi Wu, Lingting Zhu, Lei Liu, Wandi Qiao 외

Multimodal autoregressive (AR) models, based on next-token prediction and transformer architecture, have demonstrated remarkable capabilities in various multimodal tasks including text-to-image (T2I) generation. Despite …

Image Generation