paper-with-me

홈 › Papers

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings

2026-04-28 · Luca Parolari, Nicla Faccioli, Lamberto Ballan arxiv

Evaluating layout-guided text-to-image generative models requires assessing both semantic alignment with textual prompts and spatial fidelity to prescribed layouts. Assessing layout alignment requires collecting fine-grained annotations, which is costly and labor-intensive. Consequently, current benchmarks rarely provide comprehensive layout evaluation and often remain limited in scale or coverage, making model comparison, ranking, and interpretation difficult. In this work, we introduce a closed-set benchmark (C-Bench) designed to isolate key generative capabilities while providing varying levels of complexity in both prompt structure and layout. To complement this controlled setting, we propose an open-set benchmark (O-Bench) that evaluates models using real-world prompts and layouts, offering a measure of semantic and spatial alignment in the wild. We further develop a unified evaluation protocol that combines semantic and spatial accuracy into a single score, ensuring consistent model ranking. Using our benchmarks, we conduct a large-scale evaluation of six state-of-the-art layout-guided diffusion models, totaling 319,086 generated and evaluated images. We establish a model ranking based on their overall performance and provide detailed breakdowns for text and layout alignment to enhance interpretability. Fine-grained analyses across scenarios and prompt complexities highlight the strengths and limitations of current models. Code is available at https://github.com/lparolari/cobench.

📄 PDF Abstract BibTeX arXiv:2604.25358

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing

2024-10-03 · Kaizhi Zheng, Xiaotong Chen, Xuehai He, Jing Gu 외

Given the steep learning curve of professional 3D software and the time-consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in fields such as virtual reality, augment…

3D scene Editing

Layout-Guided Controllable Pathology Image Generation with In-Context Diffusion Transformers

2026-03-11 · Yuntao Shou, Xiangyong Cao, Qian Zhao, Deyu Meng arxiv

Controllable pathology image synthesis requires reliable regulation of spatial layout, tissue morphology, and semantic detail. However, existing text-guided diffusion models offer only coarse global control and lack the …

Cancer ClassificationData AugmentationImage Generation

PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models

2025-03-13 · Runze He, Bo Cheng, Yuhang Ma, Qingxiang Jia 외

In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images. Unlike previous diffusion-based models that treat layout pla…

Image GenerationImage ManipulationLayout-to-Image Generation

POCI-Diff: Position Objects Consistently and Interactively with 3D-Layout Guided Diffusion

2026-01-20 · Andrea Rigo, Luca Stornaiuolo, Weijie Wang, Mauro Martino 외 arxiv

We propose a diffusion-based approach for Text-to-Image (T2I) generation with consistent and interactive 3D layout control and editing. While prior methods improve spatial adherence using 2D cues or iterative copy-warp-p…

ConsistCompose: Unified Multimodal Layout Control for Image Composition

2025-11-23 · Xuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 외 arxiv

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counter…

Visual GroundingImage Generation