paper-with-me

홈 › Papers

UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations

2026-03-11 · Yaqi Zhao, Wang Lin, Zijian Zhang, Miles Yang, Jingyuan Chen, Wentao Zhang, Zhao Zhong, Liefeng Bo arxiv

Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual understanding tasks. Conversely, directly modeling continuous semantic representations (e.g., CLIP, SigLIP) poses significant challenges in high-dimensional generative modeling, resulting in slow convergence and training instability. To resolve this dilemma, we introduce UniCom, a unified framework that harmonizes multimodal understanding and generation via compressed continuous representation. We empirically demonstrate that reducing channel dimension is significantly more effective than spatial downsampling for both reconstruction and generation. Accordingly, we design an attention-based semantic compressor to distill dense features into a compact unified representation. Furthermore, we validate that the transfusion architecture surpasses query-based designs in convergence and consistency. Experiments demonstrate that UniCom achieves state-of-the-art generation performance among unified models. Notably, by preserving rich semantic priors, it delivers exceptional controllability in image editing and maintains image consistency even without relying on VAE.

📄 PDF Abstract BibTeX arXiv:2603.10702

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization, and Distillation

2026-02-09 · Jonathan von Rad, Yong Cao, Andreas Geiger arxiv

Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning and quantization evaluated primarily on knowledge-centric benchmarks. Thus…

Knowledge DistillationModel Compression

UniCompress: Token Compression for Unified Vision-Language Understanding and Generation

2026-03-11 · Ziyao Wang, Chen Chen, Jingtao Li, Weiming Zhuang 외 arxiv

Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework. This unified design offers architectura…

Pixel-Space Diffusion Transformers

2026-07-20 · Renye Yan, Jikang Cheng, You Wu, Ling Liang 외 arxiv

Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while sepa…

Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion

2026-01-07 · Yuanfeng Xu, Yuhao Chen, Liang Lin, Guangrun Wang arxiv

The bifurcation of generative modeling into autoregressive approaches for discrete data (text) and diffusion approaches for continuous data (images) hinders the development of truly unified multimodal systems. While Mask…

multimodal generationImage Generation

Multi-Agent Reinforcement Learning for Traffic Signal Control through Universal Communication Method

2022-04-26 · Qize Jiang, Minhao Qin, Shengmin Shi, Weiwei Sun 외

How to coordinate the communication among intersections effectively in real complex traffic scenarios with multi-intersection is challenging. Existing approaches only enable the communication in a heuristic manner withou…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning (RL)Traffic Signal Control