paper-with-me

홈 › Papers

Customization Assistant for Text-to-image Generation

2023-12-05 · CVPR 2024 1 · Yufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Tong Sun

Customizing pre-trained text-to-image generation model has attracted massive research interest recently, due to its huge potential in real-world applications. Although existing methods are able to generate creative content for a novel concept contained in single user-input image, their capability are still far from perfection. Specifically, most existing methods require fine-tuning the generative model on testing images. Some existing methods do not require fine-tuning, while their performance are unsatisfactory. Furthermore, the interaction between users and models are still limited to directive and descriptive prompts such as instructions and captions. In this work, we build a customization assistant based on pre-trained large language model and diffusion model, which can not only perform customized generation in a tuning-free manner, but also enable more user-friendly interactions: users can chat with the assistant and input either ambiguous text or clear instruction. Specifically, we propose a new framework consists of a new model design and a novel training strategy. The resulting assistant can perform customized generation in 2-5 seconds without any test time fine-tuning. Extensive experiments are conducted, competitive results have been obtained across different domains, illustrating the effectiveness of the proposed method.

📄 PDF Abstract BibTeX arXiv:2312.03045

Code (1)

drboog/profusion jax

Tasks

DescriptiveImage GenerationLanguage ModellingLarge Language ModelText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

User-Friendly Customized Generation with Multi-Modal Prompts

2024-05-26 · Linhao Zhong, Yan Hong, Wentao Chen, Binglin Zhou 외

Text-to-image generation models have seen considerable advancement, catering to the increasing interest in personalized image creation. Current customization techniques often necessitate users to provide multiple images …

DescriptiveImage GenerationText to Image GenerationText-to-Image Generation

GroundingBooth: Grounding Text-to-Image Customization

2024-09-13 · Zhexiao Xiong, Wei Xiong, Jing Shi, He Zhang 외

Recent studies in text-to-image customization show great success in generating personalized object variants given several images of a subject. While existing methods focus more on preserving the identity of the subject, …

Image Generation

Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Adapter

2024-06-05 · Peng Xing, Ning Wang, Jianbo Ouyang, Zechao Li

The remarkable advancement in text-to-image generation models significantly boosts the research in ID customization generation. However, existing personalization methods cannot simultaneously satisfy high fidelity and hi…

Image GenerationText to Image GenerationText-to-Image Generation

Real-time ASR Customization via Hypotheses Re-ordering: A Comparative Study of Different Scoring Functions

2022-01-16 · ACL ARR January 2022 1 · Anonymous

General purpose automatic speech recognizers (ASRs) require customization to the domain and context, to achieve practically acceptable accuracy levels when used as part of voice digital assistants. Further, such general …

Multi-party Collaborative Attention Control for Image Customization

2025-01-01 · CVPR 2025 1 · Han Yang, Chuanguang Yang, Qiuli Wang, Zhulin An 외

The rapid development of diffusion models has fueled a growing demand for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditi…

Image Generation