paper-with-me

홈 › Papers

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

2024-01-01 · Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin, JianFeng Wang, Kevin Lin, Zhengyuan Yang, Lijuan Wang, Mike Zheng Shou

In the evolution of Vision-Language Pre-training, shifting from short-text comprehension to encompassing extended textual contexts is pivotal. Recent autoregressive vision-language models like \cite{flamingo, palme}, leveraging the long-context capability of Large Language Models, have excelled in few-shot text generation tasks but face challenges in alignment tasks. Addressing this gap, we introduce the contrastive loss into text generation models, presenting the COntrastive-Streamlined MultimOdal framework (\ModelName), strategically partitioning the language model into dedicated unimodal text processing and adept multimodal data handling components. \ModelName, our unified framework, merges unimodal and multimodal elements, enhancing model performance for tasks involving textual and visual data while notably reducing learnable parameters. However, these models demand extensive long-text datasets, yet the availability of high-quality long-text video datasets remains limited. To bridge this gap, this work introduces \VideoDatasetName, an inaugural interleaved video-text dataset featuring comprehensive captions, marking a significant step forward. Demonstrating its impact, we illustrate how \VideoDatasetName{} enhances model performance in image-text tasks. With 34% learnable parameters and utilizing 72\% of the available data, our model demonstrates significant superiority over OpenFlamingo~\cite{openflamingo}. For instance, in the 4-shot flickr captioning task, performance notably improves from 57.2% to 65.\%. The contributions of \ModelName{} and \VideoDatasetName{} are underscored by notable performance gains across 14 diverse downstream datasets encompassing both image-text and video-text tasks.

📄 PDF Abstract BibTeX arXiv:2401.00849

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingReading ComprehensionText Generation

Similar Papers 제목 키워드 기반

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

2026-04-16 · Ziyang Luo, Nian Liu, Junwei Han arxiv

Omni-modal Large Language Models (Omni-MLLMs) promise a unified integration of diverse sensory streams. However, recent evaluations reveal a critical performance paradox: unimodal baselines frequently outperform joint mu…

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings

2026-04-21 · Zijie Li, Yichun Shi, Jingxiang Sun, Ye Wang 외 arxiv

We present MMCORE, a unified framework designed for multimodal image generation and editing. MMCORE leverages a pre-trained Vision-Language Model (VLM) to predict semantic visual embeddings via learnable query tokens, wh…

Spatial ReasoningVisual GroundingImage GenerationImage Editing

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

2026-06-29 · Chonghuinan Wang, Zhikai Chen, Chunwei Wang, Yecong Wan 외 arxiv

The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving the interleaving of both modalities. To …

Image GenerationStyle Transfer

Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization

2026-03-10 · Ming Nie, Chunwei Wang, Jianhua Han, Hang Xu 외 arxiv

Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks l…

Text-to-Image GenerationReinforcement LearningVisual StorytellingVisual Reasoning

RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm

2025-02-18 · Tiancheng Gu, Kaicheng Yang, Chaoyi Zhang, Yin Xie 외

After pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved…

Representation LearningRetrieval