Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation
Time consumption and the complexity of manual layout design make automated layout generation a critical task, especially for multiple applications across different mobile devices. Existing graph-based layout generation approaches suffer from limited generative capability, often resulting in unreasonable and incompatible outputs. Meanwhile, vision based generative models tend to overlook the original structural information, leading to component intersections and overlaps. To address these challenges, we propose an Aggregation Structural Representation (ASR) module that integrates graph networks with large language models (LLMs) to preserve structural information while enhancing generative capability. This novel pipeline utilizes graph features as hierarchical prior knowledge, replacing the traditional Vision Transformer (ViT) module in multimodal large language models (MLLM) to predict full layout information for the first time. Moreover, the intermediate graph matrix used as input for the LLM is human editable, enabling progressive, human centric design generation. A comprehensive evaluation on the RICO dataset demonstrates the strong performance of ASR, both quantitatively using mean Intersection over Union (mIoU), and qualitatively through a crowdsourced user study. Additionally, sampling on relational features ensures diverse layout generation, further enhancing the adaptability and creativity of the proposed approach.
Code (0)
등록된 구현이 없습니다.
Tasks
Layout DesignLayout GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Differentially Categorized Structural Connectome Hubs are Involved in Differential Microstructural Basis and Functional Implications and Contribute to Individual Identification
Human brain structural networks contain sets of centrally embedded hub regions that enable efficient information communication. However, it remains largely unknown about categories of structural brain hubs and their micr…
Diffusion MRIStructural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models
Abstract grammatical knowledge - of parts of speech and grammatical patterns - is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowledge in large language models? In the hu…
SentenceEfficient Human Pose Estimation by Learning Deeply Aggregated Representations
In this paper, we propose an efficient human pose estimation network (DANet) by learning deeply aggregated representations. Most existing models explore multi-scale information mainly from features with different spatial…
CPUPose EstimationStructural Guidance for Transformer Language Models
Transformer-based language models pre-trained on large amounts of text data have proven remarkably successful in learning generic transferable linguistic representations. Here we study whether structural guidance leads t…
Language ModelingLanguage ModellingStructural Similarities Between Language Models and Neural Response Measurements
Large language models (LLMs) have complicated internal dynamics, but induce representations of words and phrases whose geometry we can study. Human language processing is also opaque, but neural response measurements can…
Brain Decoding