paper-with-me

홈 › Papers

VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling

2024-08-02 · Qian Zhang, Xiangzi Dai, Ninghua Yang, Xiang An, Ziyong Feng, Xingyu Ren

VAR is a new generation paradigm that employs 'next-scale prediction' as opposed to 'next-token prediction'. This innovative transformation enables auto-regressive (AR) transformers to rapidly learn visual distributions and achieve robust generalization. However, the original VAR model is constrained to class-conditioned synthesis, relying solely on textual captions for guidance. In this paper, we introduce VAR-CLIP, a novel text-to-image model that integrates Visual Auto-Regressive techniques with the capabilities of CLIP. The VAR-CLIP framework encodes captions into text embeddings, which are then utilized as textual conditions for image generation. To facilitate training on extensive datasets, such as ImageNet, we have constructed a substantial image-text dataset leveraging BLIP2. Furthermore, we delve into the significance of word positioning within CLIP for the purpose of caption guidance. Extensive experiments confirm VAR-CLIP's proficiency in generating fantasy images with high fidelity, textual congruence, and aesthetic excellence. Our project page are https://github.com/daixiangzi/VAR-CLIP

📄 PDF Abstract BibTeX arXiv:2408.01181

Code (1)

daixiangzi/var-clip 공식 구현 pytorch

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis

2023-01-30 · CVPR 2023 1 · Ming Tao, Bing-Kun Bao, Hao Tang, Changsheng Xu

Synthesizing high-fidelity complex images from text is challenging. Based on large pretraining, the autoregressive and diffusion models can synthesize photo-realistic images. Although these large models have shown notabl…

Image GenerationScene UnderstandingText-to-Image Generation

CLIP-CLOP: CLIP-Guided Collage and Photomontage

2022-05-06 · Piotr Mirowski, Dylan Banarse, Mateusz Malinowski, Simon Osindero 외

The unabated mystique of large-scale neural networks, such as the CLIP dual image-and-text encoder, popularized automatically generated art. Increasingly more sophisticated generators enhanced the artworks' realism and v…

Prompt Engineering

CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP

2022-03-01 · ZiHao Wang, Wei Liu, Qian He, Xinglong Wu 외

Training a text-to-image generator in the general domain (e.g., Dall.e, CogView) requires huge amounts of paired text-image data, which is too expensive to collect. In this paper, we propose a self-supervised scheme name…

Image GenerationText to Image GenerationText-to-Image Generation

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

2025-05-30 · Yinqi Li, Jiahe Zhao, Hong Chang, Ruibing Hou 외

Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP falls short in distinguishing detailed di…

Large Language ModelMultimodal Large Language Model

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

2024-11-04 · Maitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim 외

Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the eff…

Diversityimage-classificationImage ClassificationImage Retrieval+2