paper-with-me

홈 › Papers

Multitwine: Multi-Object Compositing with Text and Layout Control

2025-02-07 · CVPR 2025 1 · Gemma Canet Tarrés, Zhe Lin, Zhifei Zhang, He Zhang, Andrew Gilbert, John Collomosse, Soo Ye Kim

We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex actions requiring reposing (e.g., hugging, playing guitar). When an interaction implies additional props, like `taking a selfie', our model autonomously generates these supporting objects. By jointly training for compositing and subject-driven generation, also known as customization, we achieve a more balanced integration of textual and visual inputs for text-driven object compositing. As a result, we obtain a versatile model with state-of-the-art performance in both tasks. We further present a data generation pipeline leveraging visual and language models to effortlessly synthesize multimodal, aligned training data.

📄 PDF Abstract BibTeX arXiv:2502.05165

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories

2026-01-30 · Gemma Canet Tarrés, Manel Baradad, Francesc Moreno-Noguer, Yumeng Li arxiv

Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simultaneous (i) near-perfect preservation of e…

LayoutBERT: Masked Language Layout Model for Object Insertion

2022-04-30 · Kerem Turgutlu, Sanat Sharma, Jayant Kumar

Image compositing is one of the most fundamental steps in creative workflows. It involves taking objects/parts of several images to create a new image, called a composite. Currently, this process is done manually by crea…

Language ModellingmodelObjectRetrieval

Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation

2024-10-01 · Yunnan Wang, Ziqiang Li, Zequn Zhang, Wenyao Zhang 외

There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple object…

DisentanglementImage Generation

Towards Design Compositing

2026-04-16 · Abhinav Mahajan, Abhikhya Tripathy, Sudeeksha Reddy Pala, Vaibhav Methi 외 arxiv

Graphic design creation involves harmoniously assembling multimodal components such as images, text, logos, and other visual assets collected from diverse sources, into a visually-appealing and cohesive design. Recent me…

Thinking Outside the BBox: Unconstrained Generative Object Compositing

2024-09-06 · Gemma Canet Tarrés, Zhe Lin, Zhifei Zhang, Jianming Zhang 외

Compositing an object into an image involves multiple non-trivial sub-tasks such as object placement and scaling, color/lighting harmonization, viewpoint/geometry adjustment, and shadow/reflection generation. Recent gene…

Object