paper-with-me

Papers

Transferability Between Understanding and Generation in Unified Multimodal Models

2026-07-05 · Jiwon Kang, Heeji Yoon, Jaewoo Jung, Jaewon Min, Minkyeong Jeon, Biyeon Hwang, Sangwon Jung, Seungryong Kim arxiv

Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: whether training a capability on one task improves the same capability on the other without explicit supervision. Through controlled experiments, we empirically find that transferability depends on architecture-models with fully shared transformer backbone and a unified visual encoder exhibit consistent cross-task transfer, while loosely coupled designs show little or none. Leveraging this transferability, we propose a practical training strategy. The most straightforward way to improve a target generative capability (e.g., counting) is to fine-tune generation directly, but this can degrade visual quality due to distribution shift. Instead, we train the corresponding understanding task and let it transfer into generation, which improves capability-specific generative performance while minimizing distribution shift. We validate this across three capabilities-counting, spatial relation, and text recognition/generation-showing that cross-task transferability can be systematically exploited in UMMs.

📄 PDF Abstract BibTeX arXiv:2607.04423

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

2025-03-17 · Lijie Fan, Luming Tang, Siyang Qin, Tianhong Li 외

We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture processes multimodal image and text input…

Image CaptioningImage GenerationQuestion Answering

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

2024-10-17 · CVPR 2025 1 · Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma 외

In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due …

Visual Question Answering

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

2026-06-25 · Jinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang Jiang arxiv

Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically assess understanding and generation capabili…

EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture

2025-12-04 · Xin He, Longhui Wei, Jianbo Ouyang, Minghui Liao 외 arxiv

We propose EMMA, an efficient and unified architecture for multimodal understanding, generation and editing. Specifically, EMMA primarily consists of 1) An efficient autoencoder with a 32x compression ratio, which signif…

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

2025-06-12 · Zhiyang Xu, Jiuhai Chen, Zhaojiang Lin, Xichen Pan 외

Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite these gains, unified models often underperfo…

Image Generationmultimodal generation