paper-with-me

홈 › Papers

MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

2024-01-18 · Changyao Tian, Xizhou Zhu, Yuwen Xiong, Weiyun Wang, Zhe Chen, Wenhai Wang, Yuntao Chen, Lewei Lu, Tong Lu, Jie zhou, Hongsheng Li, Yu Qiao, Jifeng Dai

Developing generative models for interleaved image-text data has both research and practical value. It requires models to understand the interleaved sequences and subsequently generate images and text. However, existing attempts are limited by the issue that the fixed number of visual tokens cannot efficiently capture image details, which is particularly problematic in the multi-image scenarios. To address this, this paper presents MM-Interleaved, an end-to-end generative model for interleaved image-text data. It introduces a multi-scale and multi-image feature synchronizer module, allowing direct access to fine-grained image features in the previous context during the generation process. MM-Interleaved is end-to-end pre-trained on both paired and interleaved image-text corpora. It is further enhanced through a supervised fine-tuning phase, wherein the model improves its ability to follow complex multi-modal instructions. Experiments demonstrate the versatility of MM-Interleaved in recognizing visual details following multi-modal instructions and generating consistent images following both textual and visual conditions. Code and models are available at \url{https://github.com/OpenGVLab/MM-Interleaved}.

📄 PDF Abstract BibTeX arXiv:2401.10208

Code (1)

opengvlab/mm-interleaved 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

2026-06-29 · Chonghuinan Wang, Zhikai Chen, Chunwei Wang, Yecong Wan 외 arxiv

The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving the interleaving of both modalities. To …

Image GenerationStyle Transfer

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

2026-03-26 · Jinbo Xing, Zeyinzi Jiang, Yuxiang Tuo, Chaojie Mao 외 arxiv

Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only single-modality outputs. This challenge of p…

Holistic Evaluation for Interleaved Text-and-Image Generation

2024-06-20 · Minqian Liu, Zhiyang Xu, Zihao Lin, Trevor Ashby 외

Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order. Despite the emerging advancements in interleav…

Image Generation

OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation

2023-10-11 · Jie An, Zhengyuan Yang, Linjie Li, JianFeng Wang 외

This work investigates a challenging task named open-domain interleaved image-text generation, which generates interleaved texts and images following an input query. We propose a new interleaved generation framework base…

Question AnsweringText Generation

Towards Text-Image Interleaved Retrieval

2025-02-18 · Xin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang 외

Current multimodal information retrieval studies mainly focus on single-image inputs, which limits real-world applications involving multiple images and text-image interleaved content. In this work, we introduce the text…

Information RetrievalLanguage ModelingLanguage ModellingLarge Language Model+3