paper-with-me

홈 › Papers

Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control

2024-04-21 · Maria Mihaela Trusca, Wolf Nuyts, Jonathan Thomm, Robert Honig, Thomas Hofmann, Tinne Tuytelaars, Marie-Francine Moens

Current diffusion models create photorealistic images given a text prompt as input but struggle to correctly bind attributes mentioned in the text to the right objects in the image. This is evidenced by our novel image-graph alignment model called EPViT (Edge Prediction Vision Transformer) for the evaluation of image-text alignment. To alleviate the above problem, we propose focused cross-attention (FCA) that controls the visual attention maps by syntactic constraints found in the input sentence. Additionally, the syntax structure of the prompt helps to disentangle the multimodal CLIP embeddings that are commonly used in T2I generation. The resulting DisCLIP embeddings and FCA are easily integrated in state-of-the-art diffusion models without additional training of these models. We show substantial improvements in T2I generation and especially its attribute-object binding on several datasets.\footnote{Code and data will be made available upon acceptance.

📄 PDF Abstract BibTeX arXiv:2404.13766

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage GenerationSentenceText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis

2024-11-11 · Taihang Hu, Linxuan Li, Joost Van de Weijer, Hongcheng Gao 외

Although text-to-image (T2I) models exhibit remarkable generation capabilities, they frequently fail to accurately bind semantically related objects or attributes in the input prompts; a challenge termed semantic binding…

AttributeImage GenerationObject

T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation

2023-07-12 · NeurIPS 2023 11 · Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li 외

Despite the stunning ability to generate high-quality images by recent text-to-image models, current approaches often struggle to effectively compose objects with different attributes and relationships into a complex and…

AttributeImage GenerationText to Image GenerationText-to-Image Generation

VSC: Visual Search Compositional Text-to-Image Diffusion Model

2025-05-02 · Do Huu Dat, Nam Hyeonu, Po-Yuan Mao, Tae-Hyun Oh

Text-to-image diffusion models have shown impressive capabilities in generating realistic visuals from natural-language prompts, yet they often struggle with accurately binding attributes to corresponding objects, especi…

Attribute

How Bias Binds: Measuring Hidden Associations for Bias Control in Text-to-Image Compositions

2025-11-10 · Jeng-Lin Li, Ming-Ching Chang, Wei-Chao Chen arxiv

Text-to-image generative models often exhibit bias related to sensitive attributes. However, current research tends to focus narrowly on single-object prompts with limited contextual diversity. In reality, each object or…

Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models Function

2024-09-30 · Chenyi Zhuang, Ying Hu, Pan Gao

Text-to-image diffusion models particularly Stable Diffusion, have revolutionized the field of computer vision. However, the synthesis quality often deteriorates when asked to generate images that faithfully represent co…

AttributeDisentanglement