paper-with-me

Papers

LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation

2024-11-22 · Fan Deng, Yaguang Wu, Xinyang Yu, Xiangjun Huang, Jian Yang, Guangyu Yan, Qiang Xu

Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remains an area in need of further development. In this paper, we present LocRef-Diffusion, a novel, tuning-free model capable of personalized customization of multiple instances' appearance and position within an image. To enhance the precision of instance placement, we introduce a Layout-net, which controls instance generation locations by leveraging both explicit instance layout information and an instance region cross-attention module. To improve the appearance fidelity to reference images, we employ an appearance-net that extracts instance appearance features and integrates them into the diffusion model through cross-attention mechanisms. We conducted extensive experiments on the COCO and OpenImages datasets, and the results demonstrate that our proposed method achieves state-of-the-art performance in layout and appearance guided generation.

📄 PDF Abstract BibTeX arXiv:2411.15252

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization

2024-01-30 · Henglei Lv, Jiayu Xiao, Liang Li, Qingming Huang

Diffusion-based text-to-image personalization have achieved great success in generating subjects specified by users among various contexts. Even though, existing finetuning-based methods still suffer from model overfitti…

Diversity

Scale-Separated Conditioning for Style-Encoder-Free Diffusion Stylization

2026-08-20 · Jingtao Zhang, Haorui Gao, Youqing Liang, Zeming Liu arxiv

Reference-based diffusion stylization requires separating target geometry from transferable appearance. Existing tuning-based methods often rely on aligned content-style-target triplets or auxiliary visual encoders, whic…

MFTF: Mask-free Training-free Object Level Layout Control Diffusion Model

2024-12-02 · Shan Yang

Text-to-image generation models have revolutionized content creation, but diffusion-based vision-language models still face challenges in precisely controlling the shape, appearance, and positional placement of objects i…

DenoisingImage GenerationObjectText to Image Generation+1

Guide-and-Rescale: Self-Guidance Mechanism for Effective Tuning-Free Real Image Editing

2024-09-02 · Vadim Titov, Madina Khalmatova, Alexandra Ivanova, Dmitry Vetrov 외

Despite recent advances in large-scale text-to-image generative models, manipulating real images with these models remains a challenging problem. The main limitations of existing editing methods are that they either fail…

Move Anything with Layered Scene Diffusion

2024-04-10 · CVPR 2024 1 · Jiawei Ren, Mengmeng Xu, Jui-Chieh Wu, Ziwei Liu 외

Diffusion models generate images with an unprecedented level of quality, but how can we freely rearrange image layouts? Recent works generate controllable scenes via learning spatially disentangled latent codes, but thes…

DenoisingDisentanglement