Collage Diffusion
We seek to give users precise control over diffusion-based image generation by modeling complex scenes as sequences of layers, which define the desired spatial arrangement and visual attributes of objects in the scene. Collage Diffusion harmonizes the input layers to make objects fit together -- the key challenge involves minimizing changes in the positions and key visual attributes of the input layers while allowing other attributes to change in the harmonization process. We ensure that objects are generated in the correct locations by modifying text-image cross-attention with the layers' alpha masks. We preserve key visual attributes of input layers by learning specialized text representations per layer and by extending ControlNet to operate on layers. Layer input allows users to control the extent of image harmonization on a per-object basis, and users can even iteratively edit individual objects in generated images while keeping other objects fixed. By leveraging the rich information present in layer input, Collage Diffusion generates globally harmonized images that maintain desired object characteristics better than prior approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Conditional Image GenerationImage GenerationImage HarmonizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
NoiseCollage: A Layout-Aware Text-to-Image Diffusion Model Based on Noise Cropping and Merging
Layout-aware text-to-image generation is a task to generate multi-object images that reflect layout conditions in addition to text conditions. The current layout-aware text-to-image diffusion models still have several is…
DenoisingImage GenerationText to Image GenerationText-to-Image GenerationUniform spatial distribution of collagen fibril radii within tendon implies local activation of pC-collagen at individual fibrils
Collagen fibril cross-sectional radii show no systematic variation between the interior and the periphery of fibril bundles, indicating an effectively constant rate of collagen incorporation into fibrils throughout the b…
DiffCollage: Parallel Generation of Large Content with Diffusion Models
We present DiffCollage, a compositional diffusion model that can generate large content by leveraging diffusion models trained on generating pieces of the large content. Our approach is based on a factor graph representa…
Image GenerationInfinite Image GenerationMotion GenerationEvaluation of the Penetration Process of Fluorescent Collagenase Nanocapsules in a 3D Collagen Gel
One of the major limitations of nanomedicine is the scarce penetration of nanoparticles in tumoral tissues. These constrains have been tried to be solved by different strategies, such as the employ of polyethyleneglycol …
Patched Denoising Diffusion Models For High-Resolution Image Synthesis
We propose an effective denoising diffusion model for generating high-resolution images (e.g., 1024$\times$512), trained on small-size image patches (e.g., 64$\times$64). We name our algorithm Patch-DM, in which a new fe…
DenoisingImage Generation