paper-with-me

홈 › Papers

GaussianPretrain: A Simple Unified 3D Gaussian Representation for Visual Pre-training in Autonomous Driving

2024-11-19 · Shaoqing Xu, Fang Li, Shengyin Jiang, Ziying Song, Li Liu, Zhi-Xin Yang

Self-supervised learning has made substantial strides in image processing, while visual pre-training for autonomous driving is still in its infancy. Existing methods often focus on learning geometric scene information while neglecting texture or treating both aspects separately, hindering comprehensive scene understanding. In this context, we are excited to introduce GaussianPretrain, a novel pre-training paradigm that achieves a holistic understanding of the scene by uniformly integrating geometric and texture representations. Conceptualizing 3D Gaussian anchors as volumetric LiDAR points, our method learns a deepened understanding of scenes to enhance pre-training performance with detailed spatial structure and texture, achieving that 40.6% faster than NeRF-based method UniPAD with 70% GPU memory only. We demonstrate the effectiveness of GaussianPretrain across multiple 3D perception tasks, showing significant performance improvements, such as a 7.05% increase in NDS for 3D object detection, boosts mAP by 1.9% in HD map construction and 0.8% improvement on Occupancy prediction. These significant gains highlight GaussianPretrain's theoretical innovation and strong practical potential, promoting visual pre-training development for autonomous driving. Source code will be available at https://github.com/Public-BOTs/GaussianPretrain

📄 PDF Abstract BibTeX arXiv:2411.12452

Code (1)

public-bots/gaussianpretrain 공식 구현 pytorch

Tasks

3D Object DetectionAutonomous DrivingGPUNeRFobject-detectionObject DetectionScene UnderstandingSelf-Supervised Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Differentiable Ray Tracing with Gaussians for Unified Radio Propagation Simulation and View Synthesis

2026-05-08 · Niklas Vaara, Lam Huynh, Pekka Sangi, Miguel Bordallo López 외 arxiv

Explicit neural representations such as 3D Gaussian Splatting (3DGS) enable high-fidelity and real-time novel view synthesis, yet optimize for alpha-composited optical appearance rather than ray-intersectable geometry. I…

Novel View Synthesis

VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning

2022-11-21 · Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu 외

Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework t…

Audio-Visual Speech RecognitionLanguage ModellingRepresentation Learningspeech-recognition+3

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

2026-08-19 · Li-Heng Chen, Haokai Pang, Chengye Su, Jiarun Liu 외 arxiv

Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians)…

Representation LearningDynamic ReconstructionScene UnderstandingAutonomous Driving

OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection

2026-08-31 · Xiaoyan Wei, Zhimin Yao, Ruilin Yang, Wei Zhang 외 arxiv

Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but often rely on increasingly complex designs such as heavy cross-modal fu…

CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting

2025-05-28 · Kornel Howil, Joanna Waczyńska, Piotr Borycki, Tadeusz Dziarmaga 외

Gaussian Splatting (GS) has recently emerged as an efficient representation for rendering 3D scenes from 2D images and has been extended to images, videos, and dynamic 4D content. However, applying style transfer to GS-b…

Style Transfer