paper-with-me

Papers

FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

2024-05-08 · Xuehai He, Jian Zheng, Jacob Zhiyuan Fang, Robinson Piramuthu, Mohit Bansal, Vicente Ordonez, Gunnar A Sigurdsson, Nanyun Peng, Xin Eric Wang

Controllable text-to-image (T2I) diffusion models generate images conditioned on both text prompts and semantic inputs of other modalities like edge maps. Nevertheless, current controllable T2I methods commonly face challenges related to efficiency and faithfulness, especially when conditioning on multiple inputs from either the same or diverse modalities. In this paper, we propose a novel Flexible and Efficient method, FlexEControl, for controllable T2I generation. At the core of FlexEControl is a unique weight decomposition strategy, which allows for streamlined integration of various input types. This approach not only enhances the faithfulness of the generated image to the control, but also significantly reduces the computational overhead typically associated with multimodal conditioning. Our approach achieves a reduction of 41% in trainable parameters and 30% in memory usage compared with Uni-ControlNet. Moreover, it doubles data efficiency and can flexibly generate images under the guidance of multiple input conditions of various modalities.

📄 PDF Abstract BibTeX arXiv:2405.04834

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

2024-08-05 · Dongyang Liu, Shitian Zhao, Le Zhuo, Weifeng Lin 외

We present Lumina-mGPT, a family of multimodal autoregressive models capable of various vision and language tasks, particularly excelling in generating flexible photorealistic images from text descriptions. By initializi…

DecoderDepth EstimationImage GenerationQuestion Answering+3

Unified Multimodal Discrete Diffusion

2025-03-26 · Alexander Swerdlow, Mihir Prabhudesai, Siddharth Gandhi, Deepak Pathak 외

Multimodal generative models that can understand and generate across multiple modalities are dominated by autoregressive (AR) approaches, which process tokens sequentially from left to right, or top to bottom. These mode…

Image CaptioningImage GenerationQuestion AnsweringText Generation

ABC: Achieving Better Control of Multimodal Embeddings using VLMs

2025-03-01 · Benjamin Schneider, Florian Kerschbaum, Wenhu Chen

Visual embedding models excel at zero-shot tasks like visual retrieval and classification. However, these models cannot be used for tasks that contain ambiguity or require user instruction. These tasks necessitate a mult…

Image to textImage-to-Text RetrievalRetrievalText Retrieval+1

Caption Anything: Interactive Image Description with Diverse Multimodal Controls

2023-05-04 · Teng Wang, Jinrui Zhang, Junjie Fei, Hao Zheng 외

Controllable image captioning is an emerging multimodal topic that aims to describe the image with natural language following human purpose, $\textit{e.g.}$, looking at the specified regions or telling in a particular te…

controllable image captioningImage CaptioningImage DescriptionInstruction Following

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

2025-01-21 · Shiyue Zhang, Zheng Chong, Xi Lu, Wenqing Zhang 외

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a promising technique, offering the potential …

Image Generation