paper-with-me

Papers

DreamTuner: Single Image is Enough for Subject-Driven Generation

2023-12-21 · Miao Hua, Jiawei Liu, Fei Ding, Wei Liu, Jie Wu, Qian He

Diffusion-based models have demonstrated impressive capabilities for text-to-image generation and are expected for personalized applications of subject-driven generation, which require the generation of customized concepts with one or a few reference images. However, existing methods based on fine-tuning fail to balance the trade-off between subject learning and the maintenance of the generation capabilities of pretrained models. Moreover, other methods that utilize additional image encoders tend to lose important details of the subject due to encoding compression. To address these challenges, we propose DreamTurner, a novel method that injects reference information from coarse to fine to achieve subject-driven image generation more effectively. DreamTurner introduces a subject-encoder for coarse subject identity preservation, where the compressed general subject features are introduced through an attention layer before visual-text cross-attention. We then modify the self-attention layers within pretrained text-to-image models to self-subject-attention layers to refine the details of the target subject. The generated image queries detailed features from both the reference image and itself in self-subject-attention. It is worth emphasizing that self-subject-attention is an effective, elegant, and training-free method for maintaining the detailed features of customized subjects and can serve as a plug-and-play solution during inference. Finally, with additional subject-driven fine-tuning, DreamTurner achieves remarkable performance in subject-driven image generation, which can be controlled by a text or other conditions such as pose. For further details, please visit the project page at https://dreamtuner-diffusion.github.io/.

📄 PDF Abstract BibTeX arXiv:2312.13691

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Paste, Inpaint and Harmonize via Denoising: Subject-Driven Image Editing with Pre-Trained Diffusion Model

2023-06-13 · Xin Zhang, Jiaxian Guo, Paul Yoo, Yutaka Matsuo 외

Text-to-image generative models have attracted rising attention for flexible image editing via user-specified descriptions. However, text descriptions alone are not enough to elaborate the details of subjects, often comp…

DenoisingImage GenerationScene Generation

Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

2025-04-02 · Shaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng 외

Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibility. For the first challenge, moving fr…

Conditional Image GenerationImage GenerationPersonalized Image GenerationText-to-Image Generation

Make-Your-3D: Fast and Consistent Subject-Driven 3D Content Generation

2024-03-14 · Fangfu Liu, HanYang Wang, Weiliang Chen, Haowen Sun 외

Recent years have witnessed the strong power of 3D generation models, which offer a new level of creative flexibility by allowing users to guide the 3D content generation process through a single image or natural languag…

3D Generation

VideoBooth: Diffusion-based Video Generation with Image Prompts

2023-12-01 · CVPR 2024 1 · Yuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si 외

Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized conte…

Video Generation

Synthetic Depth-of-Field with a Single-Camera Mobile Phone

2018-06-11 · Neal Wadhwa, Rahul Garg, David E. Jacobs, Bryan E. Feldman 외

Shallow depth-of-field is commonly used by photographers to isolate a subject from a distracting background. However, standard cell phone cameras cannot produce such images optically, as their short focal lengths and sma…