paper-with-me

홈 › Papers

Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation

2025-10-31 · Riccardo Brioschi, Aleksandr Alekseev, Emanuele Nevali, Berkay Döner, Omar El Malki, Blagoj Mitrevski, Leandro Kieliger, Mark Collier, Andrii Maksai, Jesse Berent, Claudiu Musat, Efi Kokiopoulou arxiv

Graphic layout generation is a growing research area focusing on generating aesthetically pleasing layouts ranging from poster designs to documents. While recent research has explored ways to incorporate user constraints to guide the layout generation, these constraints often require complex specifications which reduce usability. We introduce an innovative approach exploiting user-provided sketches as intuitive constraints and we demonstrate empirically the effectiveness of this new guidance method, establishing the sketch-to-layout problem as a promising research direction, which is currently under-explored. To tackle the sketch-to-layout problem, we propose a multimodal transformer-based solution using the sketch and the content assets as inputs to produce high quality layouts. Since collecting sketch training data from human annotators to train our model is very costly, we introduce a novel and efficient method to synthetically generate training sketches at scale. We train and evaluate our model on three publicly available datasets: PubLayNet, DocLayNet and SlidesVQA, demonstrating that it outperforms state-of-the-art constraint-based methods, while offering a more intuitive design experience. In order to facilitate future sketch-to-layout research, we release O(200k) synthetically-generated sketches for the public datasets above. The datasets are available at https://github.com/google-deepmind/sketch_to_layout.

📄 PDF Abstract BibTeX arXiv:2510.27632

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DrawVideo: Generating Long Video from Storyboard Keyframe Sketches

2026-05-22 · Chuanzhi Xu, Huiqi Liang, Bang Shi, Huiming Zhang 외 arxiv

Long video generation requires high-fidelity synthesis, coherent narrative structure, and user control over extended time spans. Existing text-to-video methods often rely on a single long prompt, limiting control over po…

Video Generation

Sketch2BIM: A Multi-Agent Human-AI Collaborative Pipeline to Convert Hand-Drawn Floor Plans to 3D BIM

2025-10-16 · Abir Khan Ratul, Sanjay Acharjee, Somin Park, Md Nazmus Sakib arxiv

This study introduces a human-in-the-loop pipeline that converts unscaled, hand-drawn floor plan sketches into semantically consistent 3D BIM models. The workflow leverages multimodal large language models (MLLMs) within…

MaskSketch: Unpaired Structure-guided Masked Image Generation

2023-02-10 · CVPR 2023 1 · Dina Bashkirova, Jose Lezama, Kihyuk Sohn, Kate Saenko 외

Recent conditional image generation methods produce images of remarkable diversity, fidelity and realism. However, the majority of these methods allow conditioning only on labels or text prompts, which limits their level…

Conditional Image GenerationDiversityImage GenerationImage-to-Image Translation+2

U-Sketch: An Efficient Approach for Sketch to Image Diffusion Models

2024-03-27 · Ilias Mitsouras, Eleftherios Tsonis, Paraskevi Tzouveli, Athanasios Voulodimos

Diffusion models have demonstrated remarkable performance in text-to-image synthesis, producing realistic and high resolution images that faithfully adhere to the corresponding text-prompts. Despite their great success, …

DenoisingImage Generation

Scene Designer: a Unified Model for Scene Search and Synthesis from Sketch

2021-08-16 · Leo Sampaio Ferraz Ribeiro, Tu Bui, John Collomosse, Moacir Ponti

Scene Designer is a novel method for searching and generating images using free-hand sketches of scene compositions; i.e. drawings that describe both the appearance and relative positions of objects. Our core contributio…

Contrastive LearningGraph Neural NetworkObject