paper-with-me

홈 › Papers

LayoutBERT: Masked Language Layout Model for Object Insertion

2022-04-30 · Kerem Turgutlu, Sanat Sharma, Jayant Kumar

Image compositing is one of the most fundamental steps in creative workflows. It involves taking objects/parts of several images to create a new image, called a composite. Currently, this process is done manually by creating accurate masks of objects to be inserted and carefully blending them with the target scene or images, usually with the help of tools such as Photoshop or GIMP. While there have been several works on automatic selection of objects for creating masks, the problem of object placement within an image with the correct position, scale, and harmony remains a difficult problem with limited exploration. Automatic object insertion in images or designs is a difficult problem as it requires understanding of the scene geometry and the color harmony between objects. We propose LayoutBERT for the object insertion task. It uses a novel self-supervised masked language model objective and bidirectional multi-head self-attention. It outperforms previous layout-based likelihood models and shows favorable properties in terms of model capacity. We demonstrate the effectiveness of our approach for object insertion in the image compositing setting and other settings like documents and design templates. We further demonstrate the usefulness of the learned representations for layout-based retrieval tasks. We provide both qualitative and quantitative evaluations on datasets from diverse domains like COCO, PublayNet, and two new datasets which we call Image Layouts and Template Layouts. Image Layouts which consists of 5.8 million images with layout annotations is the largest image layout dataset to our knowledge. We also share ablation study results on the effect of dataset size, model size and class sample size for this task.

📄 PDF Abstract BibTeX arXiv:2205.00347

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingmodelObjectRetrieval

Similar Papers 제목 키워드 기반

A Continuous-Time Markov Chain Framework for Insertion Language Models

2026-06-08 · Dhruvesh Patel, Benjamin Rozonoyer, Soumitra Das, Tahira Naseem 외 arxiv

Insertion Language Models (ILMs) offer several advantages over left-to-right generation and mask-based generation. However, existing formulations of insertion-based generation have largely been ad-hoc. In this paper, we …

SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout Control

2023-12-08 · CVPR 2024 1 · Jaskirat Singh, Jianming Zhang, Qing Liu, Cameron Smith 외

The field of generative image inpainting and object insertion has made significant progress with the recent advent of latent diffusion models. Utilizing a precise object mask can greatly enhance these applications. Howev…

Image GenerationImage InpaintingLayout DesignLayout-to-Image Generation+1

LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding

2023-05-30 · Yi Tu, Ya Guo, Huan Chen, Jinyang Tang

Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-based backbones have led to significant perf…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+9

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

2022-04-18 · Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 외

Self-supervised pre-training techniques have achieved remarkable progress in Document AI. Most multimodal pre-trained models use a masked language modeling objective to learn bidirectional representations on the text mod…

cross-modal alignmentDocument AIdocument-image-classificationDocument Image Classification+16

A Two-Stage System for Layout-Controlled Image Generation using Large Language Models and Diffusion Models

2025-11-10 · Jan-Hendrik Koch, Jonas Krumme, Konrad Gadzicki arxiv

Text-to-image diffusion models exhibit remarkable generative capabilities, but lack precise control over object counts and spatial arrangements. This work introduces a two-stage system to address these compositional limi…

Image Generation