paper-with-me

홈 › Papers

ComposeAnything: Composite Object Priors for Text-to-Image Generation

2025-05-30 · Zeeshan Khan, ShiZhe Chen, Cordelia Schmid

Generating images from text involving complex and novel object arrangements remains a significant challenge for current text-to-image (T2I) models. Although prior layout-based methods improve object arrangements using spatial constraints with 2D layouts, they often struggle to capture 3D positioning and sacrifice quality and coherence. In this work, we introduce ComposeAnything, a novel framework for improving compositional image generation without retraining existing T2I models. Our approach first leverages the chain-of-thought reasoning abilities of LLMs to produce 2.5D semantic layouts from text, consisting of 2D object bounding boxes enriched with depth information and detailed captions. Based on this layout, we generate a spatial and depth aware coarse composite of objects that captures the intended composition, serving as a strong and interpretable prior that replaces stochastic noise initialization in diffusion-based T2I models. This prior guides the denoising process through object prior reinforcement and spatial-controlled denoising, enabling seamless generation of compositional objects and coherent backgrounds, while allowing refinement of inaccurate priors. ComposeAnything outperforms state-of-the-art methods on the T2I-CompBench and NSR-1K benchmarks for prompts with 2D/3D spatial arrangements, high object counts, and surreal compositions. Human evaluations further demonstrate that our model generates high-quality images with compositions that faithfully reflect the text.

📄 PDF Abstract BibTeX arXiv:2505.24086

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingImage GenerationObjectText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories

2026-01-30 · Gemma Canet Tarrés, Manel Baradad, Francesc Moreno-Noguer, Yumeng Li arxiv

Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simultaneous (i) near-perfect preservation of e…

ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion

2024-10-10 · Zitian Zhang, Frédéric Fortier-Chouinard, Mathieu Garon, Anand Bhattad 외

We present ZeroComp, an effective zero-shot 3D object compositing approach that does not require paired composite-scene images during training. Our method leverages ControlNet to condition from intrinsic images and combi…

Shadow Generation Using Diffusion Model with Geometry Prior

2025-01-01 · CVPR 2025 1 · Haonan Zhao, Qingyang Liu, Xinhao Tao, Li Niu 외

Image composition involves integrating foreground object into background image to obtain a composite image. One of the key challenges is to produce realistic shadow for the inserted foreground object. Recently, diffu…

model

Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions

2025-02-12 · Prajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul, Manish Gupta 외

Non-native speakers with limited vocabulary often struggle to name specific objects despite being able to visualize them, e.g., people outside Australia searching for numbats. Further, users may want to search for such e…

Contrastive LearningImage RetrievalRetrievalSketch-Based Image Retrieval

OPA: Object Placement Assessment Dataset

2021-07-05 · Liu Liu, Zhenchen Liu, Bo Zhang, Jiangtong Li 외

Image composition aims to generate realistic composite image by inserting an object from one image into another background image, where the placement (e.g., location, size, occlusion) of inserted object may be unreasonab…

Object