paper-with-me

Papers

SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMs

2025-01-01 · CVPR 2025 1 · Junsheng Wang, Nieqing Cao, Yan Ding, Mengying Xie, Fuqiang Gu, Chao Chen

Generating layouts from textual descriptions by large language models (LLMs) plays a crucial role in precise spatial reasoning-induced domains such as robotic object rearrangement and text-to-image generation. However, current methods face challenges in limited real-world examples, handling diverse layout descriptions and varying levels of granularity. To address these issues, a novel framework named Spatial Knowledge Enhanced Layout (SKE-Layout), is introduced. SKE-Layout integrates mixed spatial knowledge sources, leveraging both real and synthetic data to enhance spatial contexts. It utilizes diverse representations tailored to specific tasks and employs contrastive learning and multitask learning techniques for accurate spatial knowledge retrieval. This framework generates more accurate and fine-grained visual layouts for object rearrangement and text-to-image generation tasks, achieving improvements of 5%-30% compared to existing methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningImage GenerationLayout GenerationObject RearrangementSpatial ReasoningText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

ERNIE-Layout: Layout-Knowledge Enhanced Multi-modal Pre-training for Document Understanding

2022-01-16 · ACL ARR January 2022 1 · Anonymous

We propose ERNIE-Layout, a knowledge enhanced pre-training approach for visual document understanding, which incorporates layout-knowledge into the pre-training of visual document understanding to learn a better joint mu…

cross-modal alignmentDocument Classificationdocument understandingQuestion Answering

LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation

2025-09-21 · Sha Li, Stefano Petrangeli, Yu Shen, Xiang Chen 외 arxiv

While Large Language Models (LLMs) have demonstrated impressive reasoning and planning abilities in textual domains and can effectively follow instructions for complex tasks, their ability to understand and manipulate sp…

Reinforcement LearningSpatial Reasoning

La La LiDAR: Large-Scale Layout Generation from LiDAR Data

2025-08-05 · Youquan Liu, Lingdong Kong, Weidong Yang, Xin Li 외 arxiv

Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiDAR generation, they lack explicit contro…

Autonomous DrivingScene Generation

SDesc3D: Towards Layout-Aware 3D Indoor Scene Generation from Short Descriptions

2026-04-02 · Jie Feng, Jiawei Shen, Junjia Huang, Junpeng Zhang 외 arxiv

3D indoor scene generation conditioned on short textual descriptions provides a promising avenue for interactive 3D environment construction without the need for labor-intensive layout specification. Despite recent progr…

Scene Generation

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

2022-10-12 · Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo 외

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and utilization of layout-centered knowledge,…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+6