paper-with-me

홈 › Papers

Highly Personalized Text Embedding for Image Manipulation by Stable Diffusion

2023-03-15 · Inhwa Han, Serin Yang, Taesung Kwon, Jong Chul Ye

Diffusion models have shown superior performance in image generation and manipulation, but the inherent stochasticity presents challenges in preserving and manipulating image content and identity. While previous approaches like DreamBooth and Textual Inversion have proposed model or latent representation personalization to maintain the content, their reliance on multiple reference images and complex training limits their practicality. In this paper, we present a simple yet highly effective approach to personalization using highly personalized (HiPer) text embedding by decomposing the CLIP embedding space for personalization and content manipulation. Our method does not require model fine-tuning or identifiers, yet still enables manipulation of background, texture, and motion with just a single image and target text. Through experiments on diverse target texts, we demonstrate that our approach produces highly personalized and complex semantic image edits across a wide range of tasks. We believe that the novel understanding of the text embedding space presented in this work has the potential to inspire further research across various tasks.

📄 PDF Abstract BibTeX arXiv:2303.08767

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationImage Manipulation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace

2024-06-30 · Shian Du, Xiaotian Cheng, Qi Qian, Henglu Wei 외

Personalized text-to-image generation has attracted unprecedented attention in the recent few years due to its unique capability of generating highly-personalized images via using the input concept dataset and novel text…

Image GenerationRepresentation LearningText to Image GenerationText-to-Image Generation

Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting

2025-12-23 · Sangoh Lee, Sangwoo Mo, Wook-Shin Han arxiv

While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one specific instance among visually similar…

CatVersion: Concatenating Embeddings for Diffusion-Based Text-to-Image Personalization

2023-11-24 · Ruoyu Zhao, Mingrui Zhu, Shiyin Dong, Nannan Wang 외

We propose CatVersion, an inversion-based method that learns the personalized concept through a handful of examples. Subsequently, users can utilize text prompts to generate images that embody the personalized concept, t…

Image GenerationPersonalized Image Generation

ViCo: Plug-and-play Visual Condition for Personalized Text-to-image Generation

2023-06-01 · Shaozhe Hao, Kai Han, Shihao Zhao, Kwan-Yee K. Wong

Personalized text-to-image generation using diffusion models has recently emerged and garnered significant interest. This task learns a novel concept (e.g., a unique toy), illustrated in a handful of images, into a gener…

Image GenerationText to Image GenerationText-to-Image Generation

Beyond Inserting: Learning Identity Embedding for Semantic-Fidelity Personalized Diffusion Generation

2024-01-31 · Yang Li, Songlin Yang, Wei Wang, Jing Dong

Advanced diffusion-based Text-to-Image (T2I) models, such as the Stable Diffusion Model, have made significant progress in generating diverse and high-quality images using text prompts alone. However, when non-famous use…

Image GenerationPersonalized Image Generation