paper-with-me

Papers

Cross-Modal Contrastive Learning for Text-to-Image Generation

2021-01-12 · CVPR 2021 1 · Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, Yinfei Yang

The output of text-to-image synthesis systems should be coherent, clear, photo-realistic scenes with high semantic fidelity to their conditioned text descriptions. Our Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) addresses this challenge by maximizing the mutual information between image and text. It does this via multiple contrastive losses which capture inter-modality and intra-modality correspondences. XMC-GAN uses an attentional self-modulation generator, which enforces strong text-image correspondence, and a contrastive discriminator, which acts as a critic as well as a feature encoder for contrastive learning. The quality of XMC-GAN's output is a major step up from previous models, as we show on three challenging datasets. On MS-COCO, not only does XMC-GAN improve state-of-the-art FID from 24.70 to 9.33, but--more importantly--people prefer XMC-GAN by 77.3 for image quality and 74.1 for image-text alignment, compared to three other recent models. XMC-GAN also generalizes to the challenging Localized Narratives dataset (which has longer, more detailed descriptions), improving state-of-the-art FID from 48.70 to 14.12. Lastly, we train and evaluate XMC-GAN on the challenging Open Images data, establishing a strong benchmark FID score of 26.91.

📄 PDF Abstract BibTeX arXiv:2101.04702

Code (1)

google-research/xmcgan_image_generation 공식 구현 jax

Tasks

Contrastive LearningGenerative Adversarial NetworkImage GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Contrastive Parameter Disentanglement for Multi-modal Remote Sensing Image Generation

2026-07-26 · Yu Zhang, Wenda Zhao, Haojun Tang, Haipeng Wang arxiv

Existing remote sensing image generation methods are largely confined to single-modality synthesis and therefore fail to exploit the complementary information inherent in multimodal imagery. To address this limitation, w…

Image Generation

UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning

2020-12-31 · ACL 2021 5 · Wei Li, Can Gao, guocheng niu, Xinyan Xiao 외

Existed pre-training methods either focus on single-modal tasks or multi-modal tasks, and cannot effectively adapt to each other. They can only utilize single-modal data (i.e. text or image) or limited multi-modal data (…

Contrastive LearningCross-Modal RetrievalImage Captioning

Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

2024-01-16 · Yuhui Zhang, Elaine Sui, Serena Yeung-Levy

Building cross-modal applications is challenging due to limited paired multi-modal data. Recent works have shown that leveraging a pre-trained multi-modal contrastive representation space enables cross-modal tasks to be …

Image GenerationText to Image GenerationText-to-Image GenerationVideo Captioning

Cross-Modal Contrastive Representation Learning for Audio-to-Image Generation

2022-07-20 · HaeChun Chung, JooYong Shim, Jong-Kook Kim

Multiple modalities for certain information provide a variety of perspectives on that information, which can improve the understanding of the information. Thus, it may be crucial to generate data of different modality fr…

Image GenerationRepresentation Learning

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

2025-03-15 · Shentong Mo, Zehua Chen, Fan Bao, Jun Zhu

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignme…

AudioCapsAudio GenerationDenoising