paper-with-me

Papers

ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven Generation

2024-12-11 · Daniel Winter, Asaf Shul, Matan Cohen, Dana Berman, Yael Pritch, Alex Rav-Acha, Yedid Hoshen

This paper introduces a tuning-free method for both object insertion and subject-driven generation. The task involves composing an object, given multiple views, into a scene specified by either an image or text. Existing methods struggle to fully meet the task's challenging objectives: (i) seamlessly composing the object into the scene with photorealistic pose and lighting, and (ii) preserving the object's identity. We hypothesize that achieving these goals requires large scale supervision, but manually collecting sufficient data is simply too expensive. The key observation in this paper is that many mass-produced objects recur across multiple images of large unlabeled datasets, in different scenes, poses, and lighting conditions. We use this observation to create massive supervision by retrieving sets of diverse views of the same object. This powerful paired dataset enables us to train a straightforward text-to-image diffusion architecture to map the object and scene descriptions to the composited image. We compare our method, ObjectMate, with state-of-the-art methods for object insertion and subject-driven generation, using a single or multiple references. Empirically, ObjectMate achieves superior identity preservation and more photorealistic composition. Differently from many other multi-reference methods, ObjectMate does not require slow test-time tuning.

📄 PDF Abstract BibTeX arXiv:2412.08645

Code (0)

등록된 구현이 없습니다.

Tasks

Object

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation

2025-05-26 · Yu Xu, Fan Tang, You Wu, Lin Gao 외

Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly "brushes" user-specified objects into a given image guided by textual prompts. …

In-Context Learning

Magic Insert: Style-Aware Drag-and-Drop

2024-07-02 · Nataniel Ruiz, Yuanzhen Li, Neal Wadhwa, Yael Pritch 외

We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This…

Domain AdaptationObject

InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes

2024-01-10 · Mohamad Shahbazi, Liesbeth Claessens, Michael Niemeyer, Edo Collins 외

We introduce InseRF, a novel method for generative object insertion in the NeRF reconstructions of 3D scenes. Based on a user-provided textual description and a 2D bounding box in a reference viewpoint, InseRF generates …

3D scene EditingDepth EstimationMonocular Depth EstimationNeRF+2

FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors

2025-05-02 · Chenxi Li, Weijie Wang, Qiang Li, Bruno Lepri 외

Text-driven object insertion in 3D scenes is an emerging task that enables intuitive scene editing through natural language. However, existing 2D editing-based methods often rely on spatial priors such as 2D masks or 3D …

ObjectSpatial Reasoning

OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models

2025-09-22 · Jinshu Chen, Xinghui Li, Xu Bai, Tianxiang Ma 외 arxiv

Recent advances in video insertion based on diffusion models are impressive. However, existing methods rely on complex control signals but struggle with subject consistency, limiting their practical applicability. In thi…