paper-with-me

Papers

Divide & Bind Your Attention for Improved Generative Semantic Nursing

2023-07-20 · Yumeng Li, Margret Keuper, Dan Zhang, Anna Khoreva

Emerging large-scale text-to-image generative models, e.g., Stable Diffusion (SD), have exhibited overwhelming results with high fidelity. Despite the magnificent progress, current state-of-the-art models still struggle to generate images fully adhering to the input prompt. Prior work, Attend & Excite, has introduced the concept of Generative Semantic Nursing (GSN), aiming to optimize cross-attention during inference time to better incorporate the semantics. It demonstrates promising results in generating simple prompts, e.g., "a cat and a dog". However, its efficacy declines when dealing with more complex prompts, and it does not explicitly address the problem of improper attribute binding. To address the challenges posed by complex prompts or scenarios involving multiple entities and to achieve improved attribute binding, we propose Divide & Bind. We introduce two novel loss objectives for GSN: a novel attendance loss and a binding loss. Our approach stands out in its ability to faithfully synthesize desired objects with improved attribute alignment from complex prompts and exhibits superior performance across multiple evaluation benchmarks.

📄 PDF Abstract BibTeX arXiv:2307.10864

Code (1)

boschresearch/Divide-and-Bind 공식 구현 pytorch

Tasks

AttributeGenerative Semantic NursingText-to-Image Generation

Methods 이 논문이 사용한 방법론

Latent Diffusion Model Diffusion models applied to latent spaces, which are normally built with (Variational) Autoencoders.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory

2024-11-25 · Eric Hanchen Jiang, Yasi Zhang, Zhi Zhang, Yixin Wan 외

Text-to-image (T2I) diffusion models have revolutionized generative modeling by producing high-fidelity, diverse, and visually realistic images from textual prompts. Despite these advances, existing models struggle with …

AttributeDenoising

Create Your World: Lifelong Text-to-Image Diffusion

2023-09-08 · Gan Sun, Wenqi Liang, Jiahua Dong, Jun Li 외

Text-to-image generative models can produce diverse high-quality images of concepts with a text prompt, which have demonstrated excellent ability in image generation, image translation, etc. We in this work study the pro…

AttributeImage Generation

Your Local GAN: Designing Two Dimensional Local Attention Mechanisms for Generative Models

2019-11-27 · CVPR 2020 6 · Giannis Daras, Augustus Odena, Han Zhang, Alexandros G. Dimakis

We introduce a new local sparse attention layer that preserves two-dimensional geometry and locality. We show that by just replacing the dense attention layer of SAGAN with our construction, we obtain very significant FI…

Conditional Image GenerationDeep AttentionImage Generation

Harmonic Self-Conditioned Flow Matching for Multi-Ligand Docking and Binding Site Design

2023-10-09 · Hannes Stärk, Bowen Jing, Regina Barzilay, Tommi Jaakkola

A significant amount of protein function requires binding small molecules, including enzymatic catalysis. As such, designing binding pockets for small molecules has several impactful applications ranging from drug synthe…

Pandora's Ballot Box: Electoral Politics of Direct Democracy

2022-08-10 · Peter Buisseret, Richard Van Weelden

We study how office-seeking parties use direct democracy to shape elections. A party with a strong electoral base can benefit from using a binding referendum to resolve issues that divide its core supporters. When refere…