paper-with-me

홈 › Papers

Representation Forcing for Bottleneck-Free Unified Multimodal Models

2026-05-29 · Yuqing Wang, Zhijie Lin, Ceyuan Yang, Yang Zhao, Fei Xiao, Hao He, Qi Zhao, Zihan Ding, Fuyun Wang, Shuai Wang, Youliang Zhang, Haoqi Fan, Xihui Liu arxiv

Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation, imposing a structural bottleneck. Naively removing it introduces a quality gap, as the model must learn both high-level structure and low-level details from raw pixels. In this paper, we propose Representation Forcing (RF), a technique that closes this gap by making representation prediction a native capability of the model. Concretely, RF forces the decoder to autoregressively predict visual representations as intermediate tokens before pixels; these tokens then stay in context to guide pixel diffusion within the same backbone. By turning representations from perception outputs into generation targets, RF eliminates the need for any external generative latent space. We find that RF benefits both understanding and generation. On image generation, our pixel-space model with RF matches state-of-the-art VAE-based unified models. On image understanding, pixel-space RF generally outperforms its VAE-based variant. Together, these results offer an effective step toward end-to-end, bottleneck-free UMMs.

📄 PDF Abstract BibTeX arXiv:2605.31604

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

GIIFT: Graph-guided Inductive Image-free Multimodal Machine Translation

2025-07-24 · Jiafeng Xiong, Yuting Zhao arxiv

Multimodal Machine Translation (MMT) has demonstrated the significant help of visual information in machine translation. However, existing MMT methods face challenges in leveraging the modality gap by enforcing rigid vis…

Multimodal Machine Translation

Le MuMo JEPA: Multi-Modal Self-Supervised Representation Learning with Learnable Fusion Tokens

2026-03-25 · Ciem Cornelissen, Sam Leroux, Pieter Simoens arxiv

Self-supervised learning has emerged as a powerful paradigm for learning visual representations without manual annotations, yet most methods still operate on a single modality and therefore miss the complementary structu…

Self-Supervised LearningRepresentation Learning

Multimodal Information Bottleneck: Learning Minimal Sufficient Unimodal and Multimodal Representations

2022-10-31 · Sijie Mai, Ying Zeng, Haifeng Hu

Learning effective joint embedding for cross-modal data has always been a focus in the field of multimodal machine learning. We argue that during multimodal fusion, the generated multimodal embedding may be redundant, an…

Emotion RecognitionMultimodal Emotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

2024-05-08 · Zehan Wang, Ziang Zhang, Xize Cheng, Rongjie Huang 외

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhan…

Synesthesia via Direct Latent Augmentation:Bypassing the Decode-Encode Loop for Cross-Modal Distillation

2026-06-06 · Cristian Sbrolli, Nicolas Michel, Matteo Matteucci, Toshihiko Yamasaki arxiv

While multimodal integration significantly improves computer vision models, deploying them incurs prohibitive inference costs and requires scarce, perfectly paired datasets. Recent methods address this data bottleneck by…

Data Augmentation