paper-with-me

Papers

SLayR: Scene Layout Generation with Rectified Flow

2024-12-06 · Cameron Braunstein, Hevra Petekkaya, Jan Eric Lenssen, Mariya Toneva, Eddy Ilg

We introduce SLayR, Scene Layout Generation with Rectified flow. State-of-the-art text-to-image models achieve impressive results. However, they generate images end-to-end, exposing no fine-grained control over the process. SLayR presents a novel transformer-based rectified flow model for layout generation over a token space that can be decoded into bounding boxes and corresponding labels, which can then be transformed into images using existing models. We show that established metrics for generated images are inconclusive for evaluating their underlying scene layout, and introduce a new benchmark suite, including a carefully designed repeatable human-evaluation procedure that assesses the plausibility and variety of generated layouts. In contrast to previous works, which perform well in either high variety or plausibility, we show that our approach performs well on both of these axes at the same time. It is also at least 5x times smaller in the number of parameters and 37% faster than the baselines. Our complete text-to-image pipeline demonstrates the added benefits of an interpretable and editable intermediate representation.

📄 PDF Abstract BibTeX arXiv:2412.05003

Code (0)

등록된 구현이 없습니다.

Tasks

Layout Generation

Similar Papers 제목 키워드 기반

FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow

2026-03-20 · Zhifei Yang, Guangyao Zhai, Keyang Lu, YuYang Yin 외 arxiv

Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object databas…

Scene Generation

NAMI: Efficient Image Generation via Progressive Rectified Flow Transformers

2025-03-12 · Yuhang Ma, Bo Cheng, Shanyuan Liu, Ao Ma 외

Flow-based transformer models for image generation have achieved state-of-the-art performance with larger model parameters, but their inference deployment cost remains high. To enhance inference performance while maintai…

Image Generation

MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation

2026-05-27 · Muyao Wang, Zeke Xie, Yanhao Chen, Lixin Xiu 외 arxiv

End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page layout design, panel rendering, page composition, and lettering. Howev…

Visual Storytelling

Constructing a 3D Town from a Single Image

2025-05-21 · Kaizhi Zheng, Ruijian Zhang, Jing Gu, Jie Yang 외

Acquiring detailed 3D scenes typically demands costly equipment, multi-view data, or labor-intensive modeling. Therefore, a lightweight alternative, generating complex 3D scenes from a single top-down image, plays an ess…

3D InpaintingImage to 3DScene Generation

Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack

2025-10-01 · Nanxiang Jiang, Zhaoxin Fan, Enhan Kang, Daiheng Gao 외 arxiv

Recent advances in text-to-image (T2I) diffusion models have enabled impressive generative capabilities, but they also raise significant safety concerns due to the potential to produce harmful or undesirable content. Whi…