paper-with-me

홈 › Papers

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

2026-09-01 · Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu hf

While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guarantee learning synergy: the two objectives may reinforce each other, compete for capacity, or merely coexist. We investigate their relationship at the representation, task, and system levels in a controlled, structurally native setting without pretrained vision priors. At the representation level, we find that each objective provides useful signal to the other: generation enriches the visual features learned for understanding, while understanding strengthens vision--language alignment for generation. However, when both objectives are forced through the same computation path, one tends to dominate. A task-decoupled architecture that specializes conflicting visual computation while preserving semantic interaction avoids this asymmetric degradation. At the task level, through three case studies, we find positive bidirectional transfer when understanding and generation tasks rely on shared knowledge. At the system level, we show that an end-to-end UMM outperforms a matched planner--executor pipeline on complex tasks that explicitly require both image understanding and generation. Together, these results show that the value of UMMs extends beyond a unified interface: appropriate specialization, shared task knowledge, and end-to-end optimization can turn coexistence into synergy.

📄 PDF Abstract BibTeX arXiv:2609.01607

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

2026-05-07 · Zeyu Liu, Zanlin Ni, Yang Yue, Cheng Da 외 arxiv

Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely decoupled understanding and generation compo…

Image Generation

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

2026-05-09 · Weiren Zhao, Yi Dong, Cheng Chen arxiv

Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generatio…

Lance: Unified Multimodal Modeling by Multi-Task Synergy

2026-05-18 · Fengyi Fu, Mengqi Huang, Shaojin Wu, Yunsheng Jiang 외 arxiv

We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text-image-dominant designs, …

Video Generation

Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding

2026-04-09 · Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong 외 arxiv

Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts. While existing paradigms like Mixture-of-Transformers (MoT) …

Image Generation

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

2025-09-29 · Yang Shi, Yuhao Dong, Yue Ding, Yuran Wang 외 arxiv

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question remains unanswered by existing benchmarks: …

Image Generation