paper-with-me

홈 › Papers

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

2025-11-18 · Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu, Zhe Gan, Yinfei Yang, Zuxuan Wu, Afshin Dehghan arxiv

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipeline to strengthen the image understanding and generation capabilities while unlocking strong image editing ability. Especially, we propose a unified Reinforcement Learning (RL) strategy that improves both image generation and image editing jointly via shared reward models. To further enhance image editing performance, we propose a light Edit Instruction Alignment stage that significantly improves the editing instruction comprehension that is essential for the success of the RL training. Experimental results show that UniGen-1.5 demonstrates competitive understanding and generation performance. Specifically, UniGen-1.5 achieves 0.89 and 4.31 overall scores on GenEval and ImgEdit that surpass the state-of-the-art models such as BAGEL and reaching performance comparable to proprietary models such as GPT-Image-1.

📄 PDF Abstract BibTeX arXiv:2511.14760

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage GenerationImage Editing

Similar Papers 제목 키워드 기반

UniGen: A Unified Framework for Textual Dataset Generation Using Large Language Models

2024-06-27 · Siyuan Wu, Yue Huang, Chujie Gao, Dongping Chen 외

Large Language Models (LLMs) such as GPT-4 and Llama3 have significantly impacted various fields by enabling high-quality synthetic data generation and reducing dependence on expensive human-generated datasets. Despite t…

AttributeBenchmarkingData AugmentationDataset Generation+3

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

2025-05-20 · Rui Tian, Mingfei Gao, Mingze Xu, Jiaming Hu 외

We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pr…

Image GenerationLanguage ModelingLanguage ModellingLarge Language Model+1

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models

2026-04-28 · Jiayi Guo, Linqing Wang, Jiangshan Wang, Yang Yue 외 arxiv

Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified capability allows UMMs to refine outputs after their initial generatio…

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection

2026-04-23 · Yanran Zhang, Wenzhao Zheng, Yifei Li, Bingyao Yu 외 arxiv

In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fields have evolved distinct architectural …

Image Generation

90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development

2025-09-30 · Runxin Yang, Yuxuan Wan, Shuqing Li, Michael R. Lyu arxiv

Developing 3D games requires specialized expertise across multiple domains, including programming, 3D modeling, and engine configuration, which limits access to millions of potential creators. Recently, researchers have …