Context-Aware Synthesis and Placement of Object Instances
Learning to insert an object instance into an image in a semantically coherent manner is a challenging and interesting problem. Solving it requires (a) determining a location to place an object in the scene and (b) determining its appearance at the location. Such an object insertion model can potentially facilitate numerous image editing and scene parsing applications. In this paper, we propose an end-to-end trainable neural network for the task of inserting an object instance mask of a specified class into the semantic label map of an image. Our network consists of two generative modules where one determines where the inserted object mask should be (i.e., location and scale) and the other determines what the object mask shape (and pose) should look like. The two modules are connected together via a spatial transformation network and jointly trained. We devise a learning procedure that leverage both supervised and unsupervised data and show our model can insert an object at diverse locations with various appearances. We conduct extensive experimental validations with comparisons to strong baselines to verify the effectiveness of the proposed network.
Code (2)
Tasks
ObjectScene ParsingSimilar Papers 제목 키워드 기반
MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
Repurposing pre-trained diffusion models has been proven to be effective for NVS. However, these methods are mostly limited to a single object; directly applying such methods to compositional multi-object scenarios yield…
DenoisingNovel View SynthesisObjectASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
Class-agnostic 3D instance segmentation tackles the challenging task of segmenting all object instances, including previously unseen ones, without semantic class reliance. Current methods struggle with generalization due…
Synthetic Data Generation3D Instance SegmentationSpatial ReasoningMachine Learning for Performance-Aware Virtual Network Function Placement
With the growing demand for data connectivity, network service providers are faced with the task of reducing their capital and operational expenses while simultaneously improving network performance and addressing the in…
BIG-bench Machine LearningToward Intelligent Scene Augmentation for Context-Aware Object Placement and Sponsor-Logo Integration
Intelligent image editing increasingly relies on advances in computer vision, multimodal reasoning, and generative modeling. While vision-language models (VLMs) and diffusion models enable guided visual manipulation, exi…
Multimodal ReasoningImage EditingDisARM: Displacement Aware Relation Module for 3D Detection
We introduce Displacement Aware Relation Module (DisARM), a novel neural network module for enhancing the performance of 3D object detection in point cloud scenes. The core idea of our method is that contextual informati…
3D Object Detectionobject-detectionObject DetectionRelation