paper-with-me

홈 › Papers

POCI-Diff: Position Objects Consistently and Interactively with 3D-Layout Guided Diffusion

2026-01-20 · Andrea Rigo, Luca Stornaiuolo, Weijie Wang, Mauro Martino, Bruno Lepri, Nicu Sebe arxiv

We propose a diffusion-based approach for Text-to-Image (T2I) generation with consistent and interactive 3D layout control and editing. While prior methods improve spatial adherence using 2D cues or iterative copy-warp-paste strategies, they often distort object geometry and fail to preserve consistency across edits. To address these limitations, we introduce a framework for Positioning Objects Consistently and Interactively (POCI-Diff), a novel formulation for jointly enforcing 3D geometric constraints and instance-level semantic binding within a unified diffusion process. Our method enables explicit per-object semantic control by binding individual text descriptions to specific 3D bounding boxes through Blended Latent Diffusion, allowing one-shot synthesis of complex multi-object scenes. We further propose a warping-free generative editing pipeline that supports object insertion, removal, and transformation via regeneration rather than pixel deformation. To preserve object identity and consistency across edits, we condition the diffusion process on reference images using IP-Adapter, enabling coherent object appearance throughout interactive 3D editing while maintaining global scene coherence. Experimental results demonstrate that POCI-Diff produces high-quality images consistent with the specified 3D layouts and edits, outperforming state-of-the-art methods in both visual fidelity and layout adherence while eliminating warping-induced geometric artifacts.

📄 PDF Abstract BibTeX arXiv:2601.14056

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Immersive Language Exploration with Object Recognition and Augmented Reality

2020-05-01 · LREC 2020 5 · Benny Platte, Anett Platte, Christian Roschke, Rico Thomanek 외

The use of Augmented Reality (AR) in teaching and learning contexts for language is still young. The ideas are endless, the concrete educational offers available emerge only gradually. Educational opportunities that were…

ObjectObject Recognition

Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos

2024-08-01 · Subin Jeon, In Cho, Minsu Kim, Woong Oh Cho 외

We propose a new framework for creating and easily manipulating 3D models of arbitrary objects using casually captured videos. Our core ingredient is a novel hierarchy deformation model, which captures motions of objects…

Direction-aware Feature-level Frequency Decomposition for Single Image Deraining

2021-06-15 · Sen Deng, Yidan Feng, Mingqiang Wei, Haoran Xie 외

We present a novel direction-aware feature-level frequency decomposition network for single image deraining. Compared with existing solutions, the proposed network has three compelling characteristics. First, unlike prev…

Rain RemovalSingle Image Deraining

Generating Interactive Worlds with Text

2019-11-20 · Angela Fan, Jack Urbanek, Pratik Ringshia, Emily Dinan 외

Procedurally generating cohesive and interesting game environments is challenging and time-consuming. In order for the relationships between the game elements to be natural, common-sense has to be encoded into arrangemen…

BIG-bench Machine LearningCommon Sense Reasoning

An Interactively Reinforced Paradigm for Joint Infrared-Visible Image Fusion and Saliency Object Detection

2023-05-17 · Di Wang, JinYuan Liu, Risheng Liu, Xin Fan

This research focuses on the discovery and localization of hidden objects in the wild and serves unmanned systems. Through empirical analysis, infrared and visible image fusion (IVIF) enables hard-to-find objects apparen…

Infrared And Visible Image Fusionobject-detectionObject DetectionSalient Object Detection