paper-with-me

Papers

MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction

2025-03-13 · Yingshuang Zou, Yikang Ding, Chuanrui Zhang, Jiazhe Guo, Bohan Li, Xiaoyang Lyu, Feiyang Tan, Xiaojuan Qi, Haoqian Wang

Recent breakthroughs in radiance fields have significantly advanced 3D scene reconstruction and novel view synthesis (NVS) in autonomous driving. Nevertheless, critical limitations persist: reconstruction-based methods exhibit substantial performance deterioration under significant viewpoint deviations from training trajectories, while generation-based techniques struggle with temporal coherence and precise scene controllability. To overcome these challenges, we present MuDG, an innovative framework that integrates Multi-modal Diffusion model with Gaussian Splatting (GS) for Urban Scene Reconstruction. MuDG leverages aggregated LiDAR point clouds with RGB and geometric priors to condition a multi-modal video diffusion model, synthesizing photorealistic RGB, depth, and semantic outputs for novel viewpoints. This synthesis pipeline enables feed-forward NVS without computationally intensive per-scene optimization, providing comprehensive supervision signals to refine 3DGS representations for rendering robustness enhancement under extreme viewpoint changes. Experiments on the Open Waymo Dataset demonstrate that MuDG outperforms existing methods in both reconstruction and synthesis quality.

📄 PDF Abstract BibTeX arXiv:2503.10604

Code (0)

등록된 구현이 없습니다.

Tasks

3DGS3D Scene ReconstructionAutonomous DrivingNovel View Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MUDGUARD: Taming Malicious Majorities in Federated Learning using Privacy-Preserving Byzantine-Robust Clustering

2022-08-22 · Rui Wang, Xingkai Wang, Huanhuan Chen, Jérémie Decouchant 외

Byzantine-robust Federated Learning (FL) aims to counter malicious clients and train an accurate global model while maintaining an extremely low attack success rate. Most existing systems, however, are only robust when m…

ClusteringFederated LearningPrivacy Preserving

Multimodal Unsupervised Domain Generalization by Retrieving Across the Modality Gap

2024-02-06 · Christopher Liao, Christian So, Theodoros Tsiligkaridis, Brian Kulis

Domain generalization (DG) is an important problem that learns a model which generalizes to unseen test domains leveraging one or more source domains, under the assumption of shared label spaces. However, most DG methods…

Domain GeneralizationQuantizationRetrievalText Augmentation

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation

2025-01-28 · Chenguo Lin, Panwang Pan, Bangbang Yang, Zeming Li 외

Recent advancements in 3D content generation from text or a single image struggle with limited high-quality 3D datasets and inconsistency from 2D multi-view generation. We introduce DiffSplat, a novel 3D generative frame…

3D Generation

Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation

2025-01-09 · Xuyi Meng, Chen Wang, Jiahui Lei, Kostas Daniilidis 외

Recent advances in 2D image generation have achieved remarkable quality,largely driven by the capacity of diffusion models and the availability of large-scale datasets. However, direct 3D generation is still constrained …

3D GenerationAttributeImage GenerationImage to 3D

ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling

2025-12-01 · Qisen Wang, Yifan Zhao, Peisen Shen, Jialu Li 외 arxiv

Although prevailing camera-controlled video generation models can produce cinematic results, lifting them directly to the generation of 3D-consistent and high-fidelity time-synchronized multi-view videos remains challeng…

Data AugmentationVideo Generation