Papers Layout-to-Image Generation
“Layout-to-Image Generation” 태그가 달린 논문 50편 · 필터 해제
OccluRank: Controllable Occlusion-Aware Layout-to-Image Generation by Adding Just an Ordinal Rank
Layout-to-image generation enables explicit spatial control through bounding-box layouts, yet bounding boxes specify only instance locations and cannot represent their occlusion order. Existing methods may rely on additi…
Layout-to-Image GenerationEnvisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation
The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I methods yield fragmented and distorted generations under few-shot atypica…
Layout-to-Image GenerationVisual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection
Unmanned aerial vehicle (UAV) based object detection is a critical but challenging task, when applied in dynamically changing scenarios with limited annotated training data. Layout-to-image generation approaches have pro…
Layout-to-Image GenerationObject DetectionEchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding
In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with accurate layouts and high fidelity to text descriptions (e.g., spatial relations…
Layout-to-Image GenerationLaytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
With the development of diffusion models, enhancing spatial controllability in text-to-image generation has become a vital challenge. As a representative task for addressing this challenge, layout-to-image generation aim…
Layout-to-Image GenerationText-to-Image GenerationTerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
Remote sensing vision tasks require extensive labeled data across multiple, interconnected domains. However, current generative data augmentation frameworks are task-isolated, i.e., each vision task requires training an …
Layout-to-Image GenerationData AugmentationOverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
Despite steady progress in layout-to-image generation, current methods still struggle with layouts containing significant overlap between bounding boxes. We identify two primary challenges: (1) large overlapping regions …
Layout-to-Image GenerationInstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
Diffusion models have demonstrated remarkable capabilities in generating high-quality images. Recent advancements in Layout-to-Image (L2I) generation have leveraged positional conditions and textual descriptions to facil…
Layout-to-Image GenerationLEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction
LEARN is a layout-aware diffusion framework designed to generate pedagogically aligned illustrations for STEM education. It leverages a curated BookCover dataset that provides narrative layouts and structured visual cues…
Layout-to-Image GenerationKnowledge GraphsPsi-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models
We introduce $\Psi$-Sampler, an SMC-based framework incorporating pCNL-based initial particle sampling for effective inference-time reward alignment with a score-based generative model. Inference-time reward alignment wi…
DenoisingImage GenerationLayout-to-Image GenerationPlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models
In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images. Unlike previous diffusion-based models that treat layout pla…
Image GenerationImage ManipulationLayout-to-Image GenerationToLo: A Two-Stage, Training-Free Layout-To-Image Generation Framework For High-Overlap Layouts
Recent training-free layout-to-image diffusion models have demonstrated remarkable performance in generating high-quality images with controllable layouts. These models follow a one-stage framework: Encouraging the model…
AttributeImage GenerationLayout-to-Image GenerationCreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
Diffusion models have been recognized for their ability to generate images that are not only visually appealing but also of high artistic quality. As a result, Layout-to-Image (L2I) generation has been proposed to levera…
Image GenerationLayout GenerationLayout-to-Image GenerationBoundary Attention Constrained Zero-Shot Layout-To-Image Generation
Recent text-to-image diffusion models excel at generating high-resolution images from text but struggle with precise control over spatial composition and object counting. To address these challenges, several studies deve…
Image GenerationLayout-to-Image GenerationObject CountingHiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout generation, where common bad cases inclu…
DisentanglementImage GenerationLayout GenerationLayout-to-Image GenerationRethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area with…
Image GenerationLayout-to-Image GenerationVideo EditingTraining-free Composite Scene Generation for Layout-to-Image Synthesis
Recent breakthroughs in text-to-image diffusion models have significantly advanced the generation of high-fidelity, photo-realistic images from textual descriptions. Yet, these models often struggle with interpreting spa…
Image GenerationLayout-to-Image GenerationScene GenerationLTOS: Layout-controllable Text-Object Synthesis via Adaptive Cross-attention Fusions
Controllable text-to-image generation synthesizes visual text and objects in images with certain conditions, which are frequently applied to emoji and poster generation. Visual text rendering and layout-to-image generati…
Image GenerationLayout-to-Image GenerationObjectText to Image Generation+1ObjBlur: A Curriculum Learning Approach With Progressive Object-Level Blurring for Improved Layout-to-Image Generation
We present ObjBlur, a novel curriculum learning approach to improve layout-to-image generation models, where the task is to produce realistic images from layouts composed of boxes and labels. Our method is based on progr…
Image GenerationLayout-to-Image GenerationDivCon: Divide and Conquer for Progressive Text-to-Image Generation
Diffusion-driven text-to-image (T2I) generation has achieved remarkable advancements. To further improve T2I models' capability in numerical and spatial reasoning, the layout is employed as an intermedium to bridge large…
Image GenerationLayout-to-Image GenerationSpatial ReasoningText to Image Generation+1