paper-with-me

홈 › Papers

MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors

2025-07-29 · Shouyi Lu, Zihan Lin, Chao Lu, Huanran Wang, Guirong Zhuo, Lianqing Zheng arxiv

Autonomous driving systems rely heavily on multimodal perception data to understand complex environments. However, the long-tailed distribution of real-world data hinders generalization, especially for rare but safety-critical vehicle categories. To address this challenge, we propose MultiEditor, a dual-branch latent diffusion framework designed to edit images and LiDAR point clouds in driving scenarios jointly. At the core of our approach is introducing 3D Gaussian Splatting (3DGS) as a structural and appearance prior for target objects. Leveraging this prior, we design a multi-level appearance control mechanism--comprising pixel-level pasting, semantic-level guidance, and multi-branch refinement--to achieve high-fidelity reconstruction across modalities. We further propose a depth-guided deformable cross-modality condition module that adaptively enables mutual guidance between modalities using 3DGS-rendered depth, significantly enhancing cross-modality consistency. Extensive experiments demonstrate that MultiEditor achieves superior performance in visual and geometric fidelity, editing controllability, and cross-modality consistency. Furthermore, generating rare-category vehicle data with MultiEditor substantially enhances the detection accuracy of perception models on underrepresented classes.

📄 PDF Abstract BibTeX arXiv:2507.21872

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingPoint Clouds

Similar Papers 제목 키워드 기반

DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes

2025-08-28 · Yajiao Xiong, Xiaoyu Zhou, Yongtao Wan, Deqing Sun 외 arxiv

We present DrivingGaussian++, an efficient and effective framework for realistic reconstructing and controllable editing of surrounding dynamic autonomous driving scenes. DrivingGaussian++ models the static background us…

Autonomous Driving

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

2025-12-19 · Yun He, Francesco Pittaluga, Ziyu Jiang, Matthias Zwicker 외 arxiv

LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represents each video as an explicit 3D scene graph, decomposing the scene into…

GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation

2026-08-17 · Tianchen Deng, Xuefeng Chen, Shuang Wu, Qu Chen 외 arxiv

Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reason…

Scene UnderstandingScene GenerationVisual Grounding

HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes

2026-04-06 · Mauricio Soroco, Francesco Pittaluga, Zaid Tasneem, Abhishek Aich 외 arxiv

Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real-world testing provides. Yet existing instruction guided image editors, trained on object-centr…

Autonomous DrivingBEV Segmentation

Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation

2026-03-13 · Yifan Zhan, Zhengqing Chen, Qingjie Wang, Zhuo He 외 arxiv

A major challenge in autonomous driving is the "long tail" of safety-critical edge cases, which often emerge from unusual combinations of common traffic elements. Synthesizing these scenarios is crucial, yet current cont…

Autonomous Driving