paper-with-me

홈 › Papers

GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

2026-07-01 · Xiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang, Linna Song, Qingliang Luo, Chufan Guo, Kuifeng Su arxiv

The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment and cross-modal information interaction in multimodal fusion perception. In this work, we break from the conventional BEV paradigm and propose a new universal framework for multi-modal fusion based on 3D Gaussian representation. This approach naturally unifies multi-modal features within a shared and continuous 3D Gaussian space, effectively preserving edge and fine texture details. To achieve this, we design a novel forward-projection-based multi-modal Gaussian initialization module and a shared cross-modal Gaussian encoder that iteratively updates Gaussian properties based on an attention mechanism. GaussianFusion is inherently a task-agnostic model, with its unified Gaussian representation naturally supporting various 3D perception tasks. Extensive experiments demonstrate the generality and robustness of GaussianFusion. On the nuScenes dataset, it outperforms the 3D object detection baseline BEVFusion by 2.6 NDS. Its variant surpasses GaussFormer on 3D semantic occupancy with 1.55 mIoU improvement while using only 30% of the Gaussians and achieving a 450% speedup.

📄 PDF Abstract BibTeX arXiv:2607.00746

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object Detection

Similar Papers 제목 키워드 기반

GaussianFusionOcc: A Seamless Sensor Fusion Approach for 3D Occupancy Prediction Using 3D Gaussians

2025-07-24 · Tomislav Pavković, Mohammad-Ali Nikouei Mahani, Johannes Niedermayer, Johannes Betz arxiv

3D semantic occupancy prediction is one of the crucial tasks of autonomous driving. It enables precise and safe interpretation and navigation in complex environments. Reliable predictions rely on effective sensor fusion,…

Autonomous Driving

GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving

2025-05-27 · Shuai Liu, Quanmin Liang, Zefeng Li, Boyang Li 외

Multi-sensor fusion is crucial for improving the performance and robustness of end-to-end autonomous driving systems. Existing methods predominantly adopt either attention-based flatten fusion or bird's eye view fusion t…

Autonomous DrivingBench2DriveNavSimSensor Fusion+1

UniGaussian: Driving Scene Reconstruction from Multiple Camera Models via Unified Gaussian Representations

2024-11-22 · Yuan Ren, Guile Wu, Runhao Li, Zheyuan Yang 외

Urban scene reconstruction is crucial for real-world autonomous driving simulators. Although existing methods have achieved photorealistic reconstruction, they mostly focus on pinhole cameras and neglect fisheye cameras.…

Autonomous DrivingScene Understanding

CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting

2025-05-28 · Kornel Howil, Joanna Waczyńska, Piotr Borycki, Tadeusz Dziarmaga 외

Gaussian Splatting (GS) has recently emerged as an efficient representation for rendering 3D scenes from 2D images and has been extended to images, videos, and dynamic 4D content. However, applying style transfer to GS-b…

Style Transfer

UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting

2025-02-25 · Haoyuan Li, Yanpeng Zhou, Tao Tang, Jifei Song 외

Recent advancements in multi-modal 3D pre-training methods have shown promising efficacy in learning joint representations of text, images, and point clouds. However, adopting point clouds as 3D representation fails to f…

3DGScross-modal alignmentRetrievalzero-shot-classification+1