User-Friendly Customized Generation with Multi-Modal Prompts
Text-to-image generation models have seen considerable advancement, catering to the increasing interest in personalized image creation. Current customization techniques often necessitate users to provide multiple images (typically 3-5) for each customized object, along with the classification of these objects and descriptive textual prompts for scenes. This paper questions whether the process can be made more user-friendly and the customization more intricate. We propose a method where users need only provide images along with text for each customization topic, and necessitates only a single image per visual concept. We introduce the concept of a ``multi-modal prompt'', a novel integration of text and images tailored to each customization concept, which simplifies user interaction and facilitates precise customization of both objects and scenes. Our proposed paradigm for customized text-to-image generation surpasses existing finetune-based methods in user-friendliness and the ability to customize complex objects with user-friendly inputs. Our code is available at $\href{https://github.com/zhongzero/Multi-Modal-Prompt}{https://github.com/zhongzero/Multi-Modal-Prompt}$.
Code (1)
Tasks
DescriptiveImage GenerationText to Image GenerationText-to-Image GenerationSimilar Papers 제목 키워드 기반
CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in v…
Motion GenerationRhythmDriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
World models have demonstrated superiority in autonomous driving, particularly in the generation of multi-view driving videos. However, significant challenges still exist in generating customized driving videos. In this …
Autonomous DrivingLanguage ModelingLanguage ModellingLarge Language Model+1VC-Agent: An Interactive Agent for Customized Video Dataset Collection
Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and time-consuming. In this work, we study the …
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this p…
Human-Domain Subject-to-VideoSingle-Domain Subject-to-VideoVideo AlignmentVideo GenerationCustomization Assistant for Text-to-image Generation
Customizing pre-trained text-to-image generation model has attracted massive research interest recently, due to its huge potential in real-world applications. Although existing methods are able to generate creative conte…
DescriptiveImage GenerationLanguage ModellingLarge Language Model+2