paper-with-me

홈 › Papers

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

2025-05-20 · Rui Tian, Mingfei Gao, Mingze Xu, Jiaming Hu, Jiasen Lu, Zuxuan Wu, Yinfei Yang, Afshin Dehghan

We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen's image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to the future research.

📄 PDF Abstract BibTeX arXiv:2505.14682

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

Similar Papers 제목 키워드 기반

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

2025-11-18 · Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu 외 arxiv

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipel…

Reinforcement LearningImage GenerationImage Editing

UniGenCoder: Merging Seq2Seq and Seq2Tree Paradigms for Unified Code Generation

2025-02-18 · Liangying Shao, Yanfu Yan, Denys Poshyvanyk, Jinsong Su

Deep learning-based code generation has completely transformed the way developers write programs today. Existing approaches to code generation have focused either on the Sequence-to-Sequence paradigm, which generates tar…

Code GenerationContrastive LearningDecoderMulti-Task Learning+1

UniGen: A Unified Framework for Textual Dataset Generation Using Large Language Models

2024-06-27 · Siyuan Wu, Yue Huang, Chujie Gao, Dongping Chen 외

Large Language Models (LLMs) such as GPT-4 and Llama3 have significantly impacted various fields by enabling high-quality synthetic data generation and reducing dependence on expensive human-generated datasets. Despite t…

AttributeBenchmarkingData AugmentationDataset Generation+3

90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development

2025-09-30 · Runxin Yang, Yuxuan Wan, Shuqing Li, Michael R. Lyu arxiv

Developing 3D games requires specialized expertise across multiple domains, including programming, 3D modeling, and engine configuration, which limits access to millions of potential creators. Recently, researchers have …

UniGen: A Unified Generative Framework for Retrieval and Question Answering with Large Language Models

2023-12-18 · Xiaoxi Li, Yujia Zhou, Zhicheng Dou

Generative information retrieval, encompassing two major tasks of Generative Document Retrieval (GDR) and Grounded Answer Generation (GAR), has gained significant attention in the area of information retrieval and natura…

Answer GenerationInformation RetrievalQuestion AnsweringRetrieval