paper-with-me

Papers

Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models

2024-04-10 · Yasi Zhang, Peiyu Yu, Ying Nian Wu

Text-to-image diffusion models have shown great success in generating high-quality text-guided images. Yet, these models may still fail to semantically align generated images with the provided text prompts, leading to problems like incorrect attribute binding and/or catastrophic object neglect. Given the pervasive object-oriented structure underlying text prompts, we introduce a novel object-conditioned Energy-Based Attention Map Alignment (EBAMA) method to address the aforementioned problems. We show that an object-centric attribute binding loss naturally emerges by approximately maximizing the log-likelihood of a $z$-parameterized energy-based model with the help of the negative sampling technique. We further propose an object-centric intensity regularizer to prevent excessive shifts of objects attention towards their attributes. Extensive qualitative and quantitative experiments, including human evaluation, on several challenging benchmarks demonstrate the superior performance of our method over previous strong counterparts. With better aligned attention maps, our approach shows great promise in further enhancing the text-controlled image editing ability of diffusion models.

📄 PDF Abstract BibTeX arXiv:2404.07389

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeObject

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

2023-07-24 · ICCV 2023 1 · Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg 외

Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality w…

RetrievalText to Video RetrievalVideo AlignmentVideo Retrieval

SpotActor: Training-Free Layout-Controlled Consistent Image Generation

2024-09-07 · Jiahao Wang, Caixia Yan, Weizhan Zhang, Haonan Lin 외

Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place e…

Image Generationobject-detectionObject Detection

Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models

2025-03-06 · Rui Jiang, Xinghe Fu, Guangcong Zheng, Teng Li 외

The rapid advancement of pretrained text-driven diffusion models has significantly enriched applications in image generation and editing. However, as the demand for personalized content editing increases, new challenges …

Image GenerationObject

Attention-based Class-Conditioned Alignment for Multi-Source Domain Adaptation of Object Detectors

2024-03-14 · Atif Belal, Akhil Meethal, Francisco Perdigon Romero, Marco Pedersoli 외

Domain adaptation methods for object detection (OD) strive to mitigate the impact of distribution shifts by promoting feature alignment across source and target domains. Multi-source domain adaptation (MSDA) allows lever…

BenchmarkingDomain AdaptationObjectobject-detection+1

FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention

2025-11-15 · Peng Zhang, Zhihui Lai, Wenting Chen, Xu Wu 외 arxiv

Medical vision-language pre-training (VLP) offers significant potential for advancing medical image understanding by leveraging paired image-report data. However, existing methods are limited by Fa}lse Negatives (FaNe) i…

Semantic SegmentationImage ClassificationObject Detection