paper-with-me

홈 › Papers

ObjectComposer: Consistent Generation of Multiple Objects Without Fine-tuning

2023-10-10 · Alec Helbling, Evan Montoya, Duen Horng Chau

Recent text-to-image generative models can generate high-fidelity images from text prompts. However, these models struggle to consistently generate the same objects in different contexts with the same appearance. Consistent object generation is important to many downstream tasks like generating comic book illustrations with consistent characters and setting. Numerous approaches attempt to solve this problem by extending the vocabulary of diffusion models through fine-tuning. However, even lightweight fine-tuning approaches can be prohibitively expensive to run at scale and in real-time. We introduce a method called ObjectComposer for generating compositions of multiple objects that resemble user-specified images. Our approach is training-free, leveraging the abilities of preexisting models. We build upon the recent BLIP-Diffusion model, which can generate images of single objects specified by reference images. ObjectComposer enables the consistent generation of compositions containing multiple specific objects simultaneously, all without modifying the weights of the underlying models.

📄 PDF Abstract BibTeX arXiv:2310.06968

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors

2024-06-02 · Ohad Rahamim, Hilit Segev, Idan Achituve, Yuval Atzmon 외

Generating 3D visual scenes is at the forefront of visual generative AI, but current 3D generation techniques struggle with generating scenes with multiple high-resolution objects. Here we introduce Lay-A-Scene, which so…

3D Generation

SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation

2026-02-26 · Vaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla 외 arxiv

We identify occlusion reasoning as a fundamental yet overlooked aspect for 3D layout-conditioned generation. It is essential for synthesizing partially occluded objects with depth-consistent geometry and scale. While exi…

Text-to-Image Generation

JRM: Joint Reconstruction Model for Multiple Objects without Alignment

2026-03-27 · Qirui Wu, Yawar Siddiqui, Duncan Frost, Samir Aroudj 외 arxiv

Object-centric reconstruction seeks to recover the 3D structure of a scene through composition of independent objects. While this independence can simplify modeling, it discards strong signals that could improve reconstr…

SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency

2024-07-24 · Yiming Xie, Chun-Han Yao, Vikram Voleti, Huaizu Jiang 외

We present Stable Video 4D (SV4D), a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for vid…

NeRFNovel View SynthesisVideo Generation

VIZOR: Viewpoint-Invariant Zero-Shot Scene Graph Generation for 3D Scene Reasoning

2026-01-31 · Vivek Madhavaram, Vartika Sengar, Arkadipta De, Charu Sharma arxiv

Scene understanding and reasoning has been a fundamental problem in 3D computer vision, requiring models to identify objects, their properties, and spatial or comparative relationships among the objects. Existing approac…

Scene Graph GenerationScene Understanding