paper-with-me

홈 › Papers

Learning to Customize Text-to-Image Diffusion In Diverse Context

2024-10-14 · Taewook Kim, Wei Chen, Qiang Qiu

Most text-to-image customization techniques fine-tune models on a small set of \emph{personal concept} images captured in minimal contexts. This often results in the model becoming overfitted to these training images and unable to generalize to new contexts in future text prompts. Existing customization methods are built on the success of effectively representing personal concepts as textual embeddings. Thus, in this work, we resort to diversifying the context of these personal concepts \emph{solely} within the textual space by simply creating a contextually rich set of text prompts, together with a widely used self-supervised learning objective. Surprisingly, this straightforward and cost-effective method significantly improves semantic alignment in the textual space, and this effect further extends to the image space, resulting in higher prompt fidelity for generated images. Additionally, our approach does not require any architectural modifications, making it highly compatible with existing text-to-image customization methods. We demonstrate the broad applicability of our approach by combining it with four different baseline methods, achieving notable CLIP score improvements.

📄 PDF Abstract BibTeX arXiv:2410.10058

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Custom-Edit: Text-Guided Image Editing with Customized Diffusion Models

2023-05-25 · Jooyoung Choi, Yunjey Choi, Yunji Kim, Junho Kim 외

Text-to-image diffusion models can generate diverse, high-fidelity images based on user-provided text prompts. Recent research has extended these models to support text-guided image editing. While text guidance is an int…

text-guided-image-editing

Orthogonal Adaptation for Modular Customization of Diffusion Models

2023-12-05 · CVPR 2024 1 · Ryan Po, Guandao Yang, Kfir Aberman, Gordon Wetzstein

Customization techniques for text-to-image models have paved the way for a wide range of previously unattainable applications, enabling the generation of specific concepts across diverse contexts and styles. While existi…

MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts

2024-10-30 · Jie Zhu, Yixiong Chen, Mingyu Ding, Ping Luo 외

Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, …

Image GenerationText to Image GenerationText-to-Image Generation

In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation

2025-05-26 · Yu Xu, Fan Tang, You Wu, Lin Gao 외

Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly "brushes" user-specified objects into a given image guided by textual prompts. …

In-Context Learning

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

2024-11-27 · CVPR 2025 1 · Shengqu Cai, Eric Chan, Yunzhi Zhang, Leonidas Guibas 외

Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images of a specific instance in novel contexts…

Image GenerationLanguage ModelingLanguage Modelling