paper-with-me

홈 › Papers

A-STAR: Test-time Attention Segregation and Retention for Text-to-image Synthesis

2023-06-26 · ICCV 2023 1 · Aishwarya Agarwal, Srikrishna Karanam, K J Joseph, Apoorv Saxena, Koustava Goswami, Balaji Vasan Srinivasan

While recent developments in text-to-image generative models have led to a suite of high-performing methods capable of producing creative imagery from free-form text, there are several limitations. By analyzing the cross-attention representations of these models, we notice two key issues. First, for text prompts that contain multiple concepts, there is a significant amount of pixel-space overlap (i.e., same spatial regions) among pairs of different concepts. This eventually leads to the model being unable to distinguish between the two concepts and one of them being ignored in the final generation. Next, while these models attempt to capture all such concepts during the beginning of denoising (e.g., first few steps) as evidenced by cross-attention maps, this knowledge is not retained by the end of denoising (e.g., last few steps). Such loss of knowledge eventually leads to inaccurate generation outputs. To address these issues, our key innovations include two test-time attention-based loss functions that substantially improve the performance of pretrained baseline text-to-image diffusion models. First, our attention segregation loss reduces the cross-attention overlap between attention maps of different concepts in the text prompt, thereby reducing the confusion/conflict among various concepts and the eventual capture of all concepts in the generated output. Next, our attention retention loss explicitly forces text-to-image diffusion models to retain cross-attention information for all concepts across all denoising time steps, thereby leading to reduced information loss and the preservation of all concepts in the generated output.

📄 PDF Abstract BibTeX arXiv:2306.14544

Code (0)

등록된 구현이 없습니다.

Tasks

AllDenoisingImage Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

RobustSGPO: Search-Space Control for Agent Harness Evolution

2026-09-09 · Zibo Zhao, Jijun Shi, Mo Zhou, Zhongyuan Wang 외 arxiv

Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unresolved. We introduce RobustSGPO, which spe…

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

2023-12-05 · Cong Wang, Jiaxi Gu, Panwen Hu, Songcen Xu 외

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided vi…

Image to Video GenerationVideo Generation

Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating

2026-04-07 · Dipan Maity, Suman Mondal, Arindam Roy arxiv

We introduce Gated-SwinRMT, a family of hybrid vision transformers that combine the shifted-window attention of the Swin Transformer with the Manhattan-distance spatial decay of Retentive Networks (RMT), augmented by inp…

Early deafness leads to re-shaping of global functional connectivity beyond the auditory cortex

2019-03-28

Early sensory deprivation such as blindness or deafness shapes brain development in multiple ways. While it is established that deprived brain areas start to be engaged in the processing of stimuli from the remaining mod…

Functional Connectivity

Consistent Segregation Metrics: Addressing Structural Variations in Global Labor Markets

2025-03-04 · Ana Kujundzic, Janneke Pieters

The Index of Dissimilarity (ID), widely utilized in economic literature as a measure of segregation, is inadequate for cross-country or time series studies due to its failure to account for structural variations across c…