paper-with-me

홈 › Papers

DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation

2024-10-02 · Jing He, Haodong Li, Yongzhe Hu, Guibao Shen, Yingjie Cai, Weichao Qiu, Ying-Cong Chen

In the realm of image generation, creating customized images from visual prompt with additional textual instruction emerges as a promising endeavor. However, existing methods, both tuning-based and tuning-free, struggle with interpreting the subject-essential attributes from the visual prompt. This leads to subject-irrelevant attributes infiltrating the generation process, ultimately compromising the personalization quality in both editability and ID preservation. In this paper, we present DisEnvisioner, a novel approach for effectively extracting and enriching the subject-essential features while filtering out -irrelevant information, enabling exceptional customization performance, in a tuning-free manner and using only a single image. Specifically, the feature of the subject and other irrelevant components are effectively separated into distinctive visual tokens, enabling a much more accurate customization. Aiming to further improving the ID consistency, we enrich the disentangled features, sculpting them into more granular representations. Experiments demonstrate the superiority of our approach over existing methods in instruction response (editability), ID consistency, inference speed, and the overall image quality, highlighting the effectiveness and efficiency of DisEnvisioner. Project page: https://disenvisioner.github.io/.

📄 PDF Abstract BibTeX arXiv:2410.02067

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control

2024-05-21 · Hong Chen, Xin Wang, YiPeng Zhang, Yuwei Zhou 외

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and at…

AttributeMotion GenerationText-to-Video GenerationVideo Generation

Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion

2025-08-11 · Minseo Kim, Minchan Kwon, Dongyeun Lee, Yunho Jeon 외 arxiv

The recent demand for customized image generation raises a need for techniques that effectively extract the common concept from small sets of images. Existing methods typically rely on additional guidance, such as text p…

Contrastive LearningImage Generation

PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

2026-04-01 · Weifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li 외 arxiv

Open-Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image-text pairs for ra…

Zero-Shot Object Detection

Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models

2024-03-17 · Ruibin Li, Ruihuang Li, Song Guo, Lei Zhang

Text-driven diffusion models have significantly advanced the image editing performance by using text prompts as inputs. One crucial step in text-driven image editing is to invert the original image into a latent noise co…

Image Generation

TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt

2024-10-16 · Jiahui Yang, Donglin Di, Baorui Ma, Xun Yang 외

In recent years, advancements in generative models have significantly expanded the capabilities of text-to-3D generation. Many approaches rely on Score Distillation Sampling (SDS) technology. However, SDS struggles to ac…

3D GenerationText to 3D