paper-with-me

홈 › Papers

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation

2026-03-03 · Chao Li, Tianhong Li, Sai Vidyaranya Nuthalapati, Hong-You Chen, Satya Narayan Shukla, Jianpeng Cheng, Yonghuan Yang, Jun Xiao, Xiangjun Fan, Aashu Singh, Dina Katabi, Shlok Kumar Mishra arxiv

Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes: contrastive alignment needs near-complete visible tokens, while masked generative modeling needs heavy corruption. We introduce DREAM, a unified framework that resolves this conflict through Masking Warmup, a schedule that shifts the center of the masking distribution over training, so low and high masking ratios coexist at every step. This co-exposure lets a single jointly-trained encoder serve both objectives. The resulting stable optimization unlocks Semantically Aligned Decoding at inference: the text encoder, trained against visual embeddings at all masking ratios, can score partially generated images and select the best trajectory with as little as 12.5% of the image decoded, improving both FID and throughput. DREAM outperforms its single-objective baselines, CLIP and FLUID: on ImageNet linear-probing (+1.1%), 5-shot transfer (+4.1%), ADE20K segmentation (+1.9%), and NYU depth estimation (+6.25%) over CLIP, and on CC12M FID (+6.2%) over FLUID while maintaining CLIP Score. Together, these gains show that text-image contrastive and generative objectives, when properly unified, are synergistic rather than competing.

📄 PDF Abstract BibTeX arXiv:2603.02667

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationContrastive LearningDepth Estimation

Similar Papers 제목 키워드 기반

Adapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners

2025-12-14 · N. K. B. M. P. K. B. Narasinghe, Uthayasanker Thayasivam arxiv

Large-scale multimodal foundation models, particularly Contrastive Captioners (CoCa), have achieved state-of-the-art results by unifying contrastive alignment with generative captioning. While zero-shot transfer capabili…

parameter-efficient fine-tuningFew-Shot Image ClassificationFew-Shot LearningData Augmentation

GRACE: Generative Representation Learning via Contrastive Policy Optimization

2025-10-06 · Jiashuo Sun, Shixuan Liu, Zhaochen Su, Xianrui Zhong 외 arxiv

Prevailing methods for training Large Language Models (LLMs) as text encoders rely on contrastive losses that treat the model as a black box function, discarding its generative and reasoning capabilities in favor of stat…

Representation Learning

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

2025-03-18 · Wei Song, Yuran Wang, Zijia Song, Yadong Li 외

The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models. A vision tokenizer trained for reconstru…

Contrastive Learning

Transformed Multi-view 3D Shape Features with Contrastive Learning

2025-10-22 · Márcus Vinícius Lobo Costa, Sherlon Almeida da Silva, Bárbara Caroline Benato, Leo Sampaio Ferraz Ribeiro 외 arxiv

This paper addresses the challenges in representation learning of 3D shape features by investigating state-of-the-art backbones paired with both contrastive supervised and self-supervised learning objectives. Computer vi…

Self-Supervised LearningRepresentation LearningContrastive Learning

X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs

2024-07-18 · Sirnam Swetha, Jinyu Yang, Tal Neiman, Mamshad Nayeem Rizve 외

Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into Large Language Models (LLMs). The prevaili…

Contrastive LearningRepresentation LearningVisual Reasoning