paper-with-me

홈 › Papers

IGD: Instructional Graphic Design with Multimodal Layer Generation

2025-07-14 · Yadong Qu, Shancheng Fang, Yuxin Wang, Xiaorui Wang, Zhineng Chen, Hongtao Xie, Yongdong Zhang arxiv

Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable graphic design files at image level with poor legibility in visual text rendering, which prevents them from achieving satisfactory and practical automated graphic design. In this paper, we propose Instructional Graphic Designer (IGD) to swiftly generate multimodal layers with editable flexibility with only natural language instructions. IGD adopts a new paradigm that leverages parametric rendering and image asset generation. First, we develop a design platform and establish a standardized format for multi-scenario design files, thus laying the foundation for scaling up data. Second, IGD utilizes the multimodal understanding and reasoning capabilities of MLLM to accomplish attribute prediction, sequencing and layout of layers. It also employs a diffusion model to generate image content for assets. By enabling end-to-end training, IGD architecturally supports scalability and extensibility in complex graphic design tasks. The superior experimental results demonstrate that IGD offers a new solution for graphic design.

📄 PDF Abstract BibTeX arXiv:2507.09910

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Show and Guide: Instructional-Plan Grounded Vision and Language Model

2024-09-27 · Diogo Glória-Silva, David Semedo, João Magalhães

Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. However, existing works on plan-following …

Language ModelingLanguage ModellingMoment Retrieval

Graphic Design with Large Multimodal Model

2024-04-22 · Yutao Cheng, Zhao Zhang, Maoke Yang, Hui Nie 외

In the field of graphic design, automating the integration of design elements into a cohesive multi-layered artwork not only boosts productivity but also paves the way for the democratization of graphic design. One exist…

Layout Generationmodel

From Elements to Design: A Layered Approach for Automatic Graphic Design Composition

2024-12-27 · CVPR 2025 1 · Jiawei Lin, Shizhao Sun, Danqing Huang, Ting Liu 외

In this work, we investigate automatic design composition from multimodal graphic elements. Although recent studies have developed various generative models for graphic design, they usually face the following limitations…

COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design

2023-11-28 · Peidong Jia, Chenxuan Li, Yuhui Yuan, Zeyu Liu 외

Graphic design, which has been evolving since the 15th century, plays a crucial role in advertising. The creation of high-quality designs demands design-oriented planning, reasoning, and layer-wise generation. Unlike the…

Image Generation

CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation

2025-06-12 · Zhao Zhang, Yutao Cheng, Dexiang Hong, Maoke Yang 외

Graphic design plays a crucial role in both commercial and personal contexts, yet creating high-quality, editable, and aesthetically pleasing graphic compositions remains a time-consuming and skill-intensive task, especi…