paper-with-me

Papers

Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation

2024-05-11 · Shengyuan Liu, Bo wang, Ye Ma, Te Yang, Xipeng Cao, Quan Chen, Han Li, Di Dong, Peng Jiang

Existing subject-driven text-to-image generation models suffer from tedious fine-tuning steps and struggle to maintain both text-image alignment and subject fidelity. For generating compositional subjects, it often encounters problems such as object missing and attribute mixing, where some subjects in the input prompt are not generated or their attributes are incorrectly combined. To address these limitations, we propose a subject-driven generation framework and introduce training-free guidance to intervene in the generative process during inference time. This approach strengthens the attention map, allowing for precise attribute binding and feature injection for each subject. Notably, our method exhibits exceptional zero-shot generation ability, especially in the challenging task of compositional generation. Furthermore, we propose a novel metric GroundingScore to evaluate subject alignment thoroughly. The obtained quantitative results serve as compelling evidence showcasing the effectiveness of our proposed method. The code will be released soon.

📄 PDF Abstract BibTeX arXiv:2405.06948

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects

2024-11-28 · CVPR 2025 1 · Weimin Qiu, Jieke Wang, Meng Tang

Diffusion models have achieved unprecedented fidelity and diversity for synthesizing image, video, 3D assets, etc. However, subject mixing is a known and unresolved issue for diffusion-based image synthesis, particularly…

Image Generation

Improving Tuning-Free Real Image Editing with Proximal Guidance

2023-06-08 · Ligong Han, Song Wen, Qi Chen, Zhixing Zhang 외

DDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifier-free guidance (CFG) scales being used …

AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization

2025-12-29 · Binhe Yu, Zhen Wang, Kexin Li, Yuqian Yuan 외 arxiv

Multi-subject customization aims to synthesize multiple user-specified subjects into a coherent image. To address issues such as subjects missing or conflicts, recent works incorporate layout guidance to provide explicit…

FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation

2025-04-22 · Zebin Yao, Lei Ren, Huixing Jiang, Chen Wei 외

Subject-driven image generation aims to synthesize novel scenes that faithfully preserve subject identity from reference images while adhering to textual guidance, yet existing methods struggle with a critical trade-off …

Image GenerationText to Image GenerationText-to-Image Generation

Classifier-free guidance in LLMs Safety

2024-12-08 · Roman Smirnov

The paper describes LLM unlearning without a retaining dataset, using the ORPO reinforcement learning method with inference enhanced by modified classifier-free guidance. Significant improvement in unlearning, without de…

reinforcement-learningReinforcement Learning