paper-with-me

홈 › Papers

SAM-GAN: Self-Attention supporting Multi-stage Generative Adversarial Networks for text-to-image synthesis

2021-02-10 · Neural Networks 2021 2 · Dunlu Peng ∗, Wuchen Yang, Cong Liu, Shuairui Lü

Synthesizing photo-realistic images based on text descriptions is a challenging task in the field of computer vision. Although generative adversarial networks have made significant breakthroughs in this task, they still face huge challenges in generating high-quality visually realistic images consistent with the semantics of text. Generally, existing text-to-image methods accomplish this task with two steps, that is, first generating an initial image with a rough outline and color, and then gradually yielding the image within high-resolution from the initial image. However, one drawback of these methods is that, if the quality of the initial image generation is not high, it is hard to generate a satisfactory high-resolution image. In this paper, we propose SAM-GAN, Self-Attention supporting Multi-stage Generative Adversarial Networks, for text-to-image synthesis. With the self-attention mechanism, the model can establish the multi-level dependence of the image and fuse the sentence- and word-level visual-semantic vectors, to improve the quality of the generated image. Furthermore, a multi-stage perceptual loss is introduced to enhance the semantic similarity between the synthesized image and the real image, thus enhancing the visual-semantic consistency between text and images. For the diversity of the generated images, a mode seeking regularization term is integrated into the model. The results of extensive experiments and ablation studies, which were conducted in the Caltech-UCSD Birds and Microsoft Common Objects in Context datasets, show that our model is superior to competitive models in text-to-image synthesis.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationSemantic SimilaritySemantic Textual SimilaritySentence

Similar Papers 제목 키워드 기반

Improved Transformer for High-Resolution GANs

2021-06-14 · NeurIPS 2021 12 · Long Zhao, Zizhao Zhang, Ting Chen, Dimitris N. Metaxas 외

Attention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making them difficult to be adopted for high-reso…

Image GenerationVocal Bursts Intensity Prediction

Generative Early Stage Ranking

2025-11-26 · Juhee Hong, Meng Liu, Shengzhi Wang, Xiaoheng Mao 외 arxiv

Large-scale recommendations commonly adopt a multi-stage cascading ranking system paradigm to balance effectiveness and efficiency. Early Stage Ranking (ESR) systems utilize the "user-item decoupling" approach, where ind…

MTSIC: Multi-stage Transformer-based GAN for Spectral Infrared Image Colorization

2025-06-21 · Tingting Liu, YuAn Liu, Jinhui Tang, Liyin Yuan 외

Thermal infrared (TIR) images, acquired through thermal radiation imaging, are unaffected by variations in lighting conditions and atmospheric haze. However, TIR images inherently lack color and texture information, limi…

ColorizationGenerative Adversarial NetworkImage Colorization

Longformer: The Long-Document Transformer

2020-04-10 · Iz Beltagy, Matthew E. Peters, Arman Cohan

Transformer-based models are unable to process long sequences due to their self-attention operation, which scales quadratically with the sequence length. To address this limitation, we introduce the Longformer with an at…

DecoderLanguage ModelingLanguage ModellingQuestion Answering+1

AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache

2025-10-29 · Dinghong Song, Yuan Feng, Yiwei Wang, Shangye Chen 외 arxiv

Large Language Models (LLMs) are widely used in generative applications such as chatting, code generation, and reasoning. However, many realworld workloads such as classification, question answering, recommendation, and …

Question AnsweringCode Generation