paper-with-me

홈 › Papers

Controllable 3D Placement of Objects with Scene-Aware Diffusion Models

2025-06-26 · Mohamed Omran, Dimitris Kalatzis, Jens Petersen, Amirhossein Habibian, Auke Wiggers

Image editing approaches have become more powerful and flexible with the advent of powerful text-conditioned generative models. However, placing objects in an environment with a precise location and orientation still remains a challenge, as this typically requires carefully crafted inpainting masks or prompts. In this work, we show that a carefully designed visual map, combined with coarse object masks, is sufficient for high quality object placement. We design a conditioning signal that resolves ambiguities, while being flexible enough to allow for changing of shapes or object orientations. By building on an inpainting model, we leave the background intact by design, in contrast to methods that model objects and background jointly. We demonstrate the effectiveness of our method in the automotive setting, where we compare different conditioning signals in novel object placement tasks. These tasks are designed to measure edit quality not only in terms of appearance, but also in terms of pose and location accuracy, including cases that require non-trivial shape changes. Lastly, we show that fine location control can be combined with appearance control to place existing objects in precise locations in a scene.

📄 PDF Abstract BibTeX arXiv:2506.21446

Code (0)

등록된 구현이 없습니다.

Tasks

Object

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

La La LiDAR: Large-Scale Layout Generation from LiDAR Data

2025-08-05 · Youquan Liu, Lingdong Kong, Weidong Yang, Xin Li 외 arxiv

Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiDAR generation, they lack explicit contro…

Autonomous DrivingScene Generation

CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation

2025-02-12 · Qinghe Wang, Yawen Luo, Xiaoyu Shi, Xu Jia 외

In this work, we present CineMaster, a novel framework for 3D-aware and controllable text-to-video generation. Our goal is to empower users with comparable controllability as professional film directors: precise placemen…

ObjectText-to-Video GenerationVideo Generation

MotionCom: Automatic and Motion-Aware Image Composition with LLM and Video Diffusion Prior

2024-09-16 · Weijing Tao, Xiaofeng Yang, Miaomiao Cui, Guosheng Lin

This work presents MotionCom, a training-free motion-aware diffusion based image composition, enabling automatic and seamless integration of target objects into new scenes with dynamically coherent results without finetu…

Image GenerationLanguage ModelingLanguage Modelling

MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models

2025-12-03 · Shaoheng Fang, Chaohui Yu, Fan Wang, Qixing Huang arxiv

We introduce MVRoom, a controllable novel view synthesis (NVS) pipeline for 3D indoor scenes that uses multi-view diffusion conditioned on a coarse 3D layout. MVRoom employs a two-stage design in which the 3D layout is u…

Novel View SynthesisScene Generation

LAW-Diffusion: Complex Scene Generation by Diffusion with Layouts

2023-08-13 · ICCV 2023 1 · BinBin Yang, Yi Luo, Ziliang Chen, Guangrun Wang 외

Thanks to the rapid development of diffusion models, unprecedented progress has been witnessed in image synthesis. Prior works mostly rely on pre-trained linguistic models, but a text is often too abstract to properly sp…

Image GenerationLayout-to-Image GenerationObjectScene Generation