paper-with-me

홈 › Papers

MObI: Multimodal Object Inpainting Using Diffusion Models

2025-01-06 · Alexandru Buburuzan, Anuj Sharma, John Redford, Puneet K. Dokania, Romain Mueller

Safety-critical applications, such as autonomous driving, require extensive multimodal data for rigorous testing. Methods based on synthetic data are gaining prominence due to the cost and complexity of gathering real-world data but require a high degree of realism and controllability in order to be useful. This paper introduces MObI, a novel framework for Multimodal Object Inpainting that leverages a diffusion model to create realistic and controllable object inpaintings across perceptual modalities, demonstrated for both camera and lidar simultaneously. Using a single reference RGB image, MObI enables objects to be seamlessly inserted into existing multimodal scenes at a 3D location specified by a bounding box, while maintaining semantic consistency and multimodal coherence. Unlike traditional inpainting methods that rely solely on edit masks, our 3D bounding box conditioning gives objects accurate spatial positioning and realistic scaling. As a result, our approach can be used to insert novel objects flexibly into multimodal scenes, providing significant advantages for testing perception models.

📄 PDF Abstract BibTeX arXiv:2501.03173

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingObject

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

2025-07-30 · Alexandru Buburuzan arxiv

Safety-critical applications, such as autonomous driving and medical image analysis, require extensive multimodal data for rigorous testing. Synthetic data methods are gaining prominence due to the cost and complexity of…

Synthetic Data GenerationAutonomous Driving

Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion Model

2023-10-11 · Shiyuan Yang, Xiaodong Chen, Jing Liao

Recently, text-to-image denoising diffusion probabilistic models (DDPMs) have demonstrated impressive image generation capabilities and have also been successfully applied to image inpainting. However, in practice, users…

DenoisingImage DenoisingImage GenerationImage Inpainting

Towards Language-Driven Video Inpainting via Multimodal Large Language Models

2024-01-18 · CVPR 2024 1 · Jianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou 외

We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of traditional video inpainting methods that …

Video Inpainting

MTV-Inpaint: Multi-Task Long Video Inpainting

2025-03-14 · Shiyuan Yang, Zheng Gu, Liang Hou, Xin Tao 외

Video inpainting involves modifying local regions within a video, ensuring spatial and temporal consistency. Most existing methods focus primarily on scene completion (i.e., filling missing regions) and lack the capabili…

Image InpaintingObjectVideo Inpainting

Improving Text-guided Object Inpainting with Semantic Pre-inpainting

2024-09-12 · Yifu Chen, Jingwen Chen, Yingwei Pan, Yehao Li 외

Recent years have witnessed the success of large text-to-image diffusion models and their remarkable potential to generate high-quality images. The further pursuit of enhancing the editability of images has sparked signi…

DenoisingObject