paper-with-me

Papers

Generating Compositional Scenes via Text-to-image RGBA Instance Generation

2024-11-16 · Alessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang, Sarah Parisot

Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layout editing ability and fine-grained control over object attributes. The concept of multi-layer generation holds great potential to address these limitations, however generating image instances concurrently to scene composition limits control over fine-grained object attributes, relative positioning in 3D space and scene manipulation abilities. In this work, we propose a novel multi-stage generation paradigm that is designed for fine-grained control, flexibility and interactivity. To ensure control over instance attributes, we devise a novel training paradigm to adapt a diffusion model to generate isolated scene components as RGBA images with transparency information. To build complex images, we employ these pre-generated instances and introduce a multi-layer composite generation process that smoothly assembles components in realistic scenes. Our experiments show that our RGBA diffusion model is capable of generating diverse and high quality instances with precise control over object attributes. Through multi-layer composition, we demonstrate that our approach allows to build and manipulate images from highly complex prompts with fine-grained control over object appearance and location, granting a higher degree of control than competing methods.

📄 PDF Abstract BibTeX arXiv:2411.10913

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectPrompt Engineering

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

OmniPSD: Layered PSD Generation with Diffusion Transformer

2025-12-10 · Cheng Liu, Yiren Song, Haofan Wang, Mike Zheng Shou arxiv

Recent advances in diffusion models have greatly improved image generation and editing, yet generating or reconstructing layered PSD files with transparent alpha channels remains highly challenging. We propose OmniPSD, a…

Image Generation

Alfie: Democratising RGBA Image Generation With No $$$

2024-08-27 · Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Rita Cucchiara

Designs and artworks are ubiquitous across various creative fields, requiring graphic design skills and dedicated software to create compositions that include many graphical elements, such as logos, icons, symbols, and a…

Image GenerationImage MattingScene GenerationVisual Storytelling

TransPixeler: Advancing Text-to-Video Generation with Transparency

2025-01-06 · CVPR 2025 1 · Luozhou Wang, Yijun Li, Zhifei Chen, Jui-Hsien Wang 외

Text-to-video generative models have made significant strides, enabling diverse applications in entertainment, advertising, and education. However, generating RGBA video, which includes alpha channels for transparency, r…

Text-to-Video GenerationVideo Generation

Structural Multiplane Image: Bridging Neural View Synthesis and 3D Reconstruction

2023-03-10 · CVPR 2023 1 · Mingfang Zhang, Jinglu Wang, Xiao Li, Yifei HUANG 외

The Multiplane Image (MPI), containing a set of fronto-parallel RGBA layers, is an effective and efficient representation for view synthesis from sparse inputs. Yet, its fixed structure limits the performance, especially…

3D Reconstruction

MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation

2024-04-03 · CVPR 2024 1 · Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang, Fei Chen 외

Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, …

Image GenerationPrompt EngineeringText to Image GenerationText-to-Image Generation