paper-with-me

Papers

UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

2024-01-24 · Wei Li, Xue Xu, Jiachen Liu, Xinyan Xiao

Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific entities or scenes. This paper presents UNIMO-G, a simple multimodal conditional diffusion framework that operates on multimodal prompts with interleaved textual and visual inputs, which demonstrates a unified ability for both text-driven and subject-driven image generation. UNIMO-G comprises two core components: a Multimodal Large Language Model (MLLM) for encoding multimodal prompts, and a conditional denoising diffusion network for generating images based on the encoded multimodal input. We leverage a two-stage training strategy to effectively train the framework: firstly pre-training on large-scale text-image pairs to develop conditional image generation capabilities, and then instruction tuning with multimodal prompts to achieve unified image generation proficiency. A well-designed data processing pipeline involving language grounding and image segmentation is employed to construct multi-modal prompts. UNIMO-G excels in both text-to-image generation and zero-shot subject-driven synthesis, and is notably effective in generating high-fidelity images from complex multimodal prompts involving multiple image entities.

📄 PDF Abstract BibTeX arXiv:2401.13388

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Image GenerationDenoisingImage GenerationImage SegmentationLanguage ModellingLarge Language ModelMultimodal Large Language ModelSemantic SegmentationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training Models

2024-07-25 · Haonan Zheng, Xinyang Deng, Wen Jiang, Wenrui Li

With Vision-Language Pre-training (VLP) models demonstrating powerful multimodal interaction capabilities, the application scenarios of neural networks are no longer confined to unimodal domains but have expanded to more…

Data Augmentationmultimodal interaction

Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces

2025-06-09 · Kevin Rojas, Yuchen Zhu, Sichen Zhu, Felix X. -F. Ye 외

Diffusion models have demonstrated remarkable performance in generating unimodal data across various tasks, including image, video, and text generation. On the contrary, the joint generation of multimodal data through di…

Image GenerationText Generation

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

2025-11-03 · Yongyuan Liang, Wei Chow, Feng Li, Ziqiao Ma 외 arxiv

Unified multimodal models (UMMs) have emerged as a powerful paradigm for seamlessly unifying text and image understanding and generation. However, prevailing evaluations treat these abilities in isolation, such that task…

Question Answering

UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation

2025-11-21 · Chi Zhang, Jiepeng Wang, Youming Wang, Yuanzhi Liang 외 arxiv

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: th…

Text-to-Image Generation

DuoGen: Towards General Purpose Interleaved Multimodal Generation

2026-01-31 · Min Shi, Xiaohui Zeng, Jiannan Huang, Yin Cui 외 arxiv

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of ex…

multimodal generationVideo GenerationImage Editing