paper-with-me

홈 › Papers

Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting

2026-03-19 · Yiren Lu, Xin Ye, Burhaneddin Yaman, Jingru Luo, Zhexiao Xiong, Liu Ren, Yu Yin arxiv

Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning for various downstream tasks, such as semantic segmentation, 3D object detection, and motion prediction. However, most existing BEV perception frameworks adopt an end-to-end training paradigm, where image features are directly transformed into the BEV space and optimized solely through downstream task supervision. This formulation treats the entire perception process as a black box, often lacking explicit 3D geometric understanding and interpretability, leading to suboptimal performance. In this paper, we claim that an explicit 3D representation matters for accurate BEV perception, and we propose Splat2BEV, a Gaussian Splatting-assisted framework for BEV tasks. Splat2BEV aims to learn BEV feature representations that are both semantically rich and geometrically precise. We first pre-train a Gaussian generator that explicitly reconstructs 3D scenes from multi-view inputs, enabling the generation of geometry-aligned feature representations. These representations are then projected into the BEV space to serve as inputs for downstream tasks. Extensive experiments on nuScenes and argoverse dataset demonstrate that Splat2BEV achieves state-of-the-art performance and validate the effectiveness of incorporating explicit 3D reconstruction into BEV perception.

📄 PDF Abstract BibTeX arXiv:2603.19193

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Segmentation3D Object DetectionAutonomous Driving3D Reconstruction

Similar Papers 제목 키워드 기반

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

2026-06-11 · Hao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang 외 arxiv

Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but stop at the visible surface, while image-to-3D models generate complete shapes that are often misaligne…

3D scene Editing

NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction

2026-03-04 · Weirong Chen, Chuanxia Zheng, Ganlin Zhang, Andrea Vedaldi 외 arxiv

We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images in a feed-forward manner. Unlike pixel-aligned methods that tie geometry to per-ray predictions, our formulati…

3D ReconstructionPoint Clouds

What Matters to You? Towards Visual Representation Alignment for Robot Learning

2023-10-11 · Ran Tian, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik 외

When operating in service of people, robots need to optimize rewards aligned with end-user preferences. Since robots will rely on raw perceptual inputs like RGB images, their rewards will inevitably use visual representa…

Zero-shot Generalization

AniPixel: Towards Animatable Pixel-Aligned Human Avatar

2023-02-07 · Jinlong Fan, Jing Zhang, Zhi Hou, DaCheng Tao

Although human reconstruction typically results in human-specific avatars, recent 3D scene reconstruction techniques utilizing pixel-aligned features show promise in generalizing to new scenes. Applying these techniques …

3D Scene Reconstruction

SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenes

2026-03-24 · Zhicheng Qiu, Jiarui Meng, Tong-an Luo, Yican Huang 외 arxiv

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling…

Dynamic ReconstructionScene Parsing