paper-with-me

홈 › Papers

BEVDiffLoc: End-to-End LiDAR Global Localization in BEV View based on Diffusion Model

2025-03-14 · Ziyue Wang, Chenghao Shi, Neng Wang, Qinghua Yu, Xieyuanli Chen, Huimin Lu

Localization is one of the core parts of modern robotics. Classic localization methods typically follow the retrieve-then-register paradigm, achieving remarkable success. Recently, the emergence of end-to-end localization approaches has offered distinct advantages, including a streamlined system architecture and the elimination of the need to store extensive map data. Although these methods have demonstrated promising results, current end-to-end localization approaches still face limitations in robustness and accuracy. Bird's-Eye-View (BEV) image is one of the most widely adopted data representations in autonomous driving. It significantly reduces data complexity while preserving spatial structure and scale consistency, making it an ideal representation for localization tasks. However, research on BEV-based end-to-end localization remains notably insufficient. To fill this gap, we propose BEVDiffLoc, a novel framework that formulates LiDAR localization as a conditional generation of poses. Leveraging the properties of BEV, we first introduce a specific data augmentation method to significantly enhance the diversity of input data. Then, the Maximum Feature Aggregation Module and Vision Transformer are employed to learn robust features while maintaining robustness against significant rotational view variations. Finally, we incorporate a diffusion model that iteratively refines the learned features to recover the absolute pose. Extensive experiments on the Oxford Radar RobotCar and NCLT datasets demonstrate that BEVDiffLoc outperforms the baseline methods. Our code is available at https://github.com/nubot-nudt/BEVDiffLoc.

📄 PDF Abstract BibTeX arXiv:2503.11372

Code (1)

nubot-nudt/bevdiffloc 공식 구현 pytorch

Tasks

Autonomous DrivingData Augmentation

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Multi-Head Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images

2026-03-17 · David Skuddis, Vincent Ress, Wei Zhang, Vincent Ofosu Nyako 외 arxiv

We present BEV-SLD, a LiDAR global localization method building on the Scene Landmark Detection (SLD) concept. Unlike scene-agnostic pipelines, our self-supervised approach leverages bird's-eye-view (BEV) images to disco…

RING#: PR-by-PE Global Localization with Roto-translation Equivariant Gram Learning

2024-08-30 · Sha Lu, Xuecheng Xu, Yuxuan Wu, Haojian Lu 외

Global localization using onboard perception sensors, such as cameras and LiDARs, is crucial in autonomous driving and robotics applications when GPS signals are unreliable. Most approaches achieve global localization by…

Autonomous DrivingPose EstimationSequential Place Recognition

S-BEVLoc: BEV-based Self-supervised Framework for Large-scale LiDAR Global Localization

2025-09-11 · Chenghao Zhang, Lun Luo, Si-Yuan Cao, Xiaokai Bai 외 arxiv

LiDAR-based global localization is an essential component of simultaneous localization and mapping (SLAM), which helps loop closure and re-localization. Current approaches rely on ground-truth poses obtained from GPS or …

Adaptive Motorized LiDAR Scanning Control for Robust Localization with OpenStreetMap

2025-09-15 · Jianping Li, Kaisong Zhu, Zhongyuan Liu, Rui Jin 외 arxiv

LiDAR-to-OpenStreetMap (OSM) localization has gained increasing attention, as OSM provides lightweight global priors such as building footprints. These priors enhance global consistency for robot navigation, but OSM is o…

Robot Navigation

Diffusion Based Robust LiDAR Place Recognition

2025-04-16 · Benjamin Krummenacher, Jonas Frey, Turcan Tuna, Olga Vysotska 외

Mobile robots on construction sites require accurate pose estimation to perform autonomous surveying and inspection missions. Localization in construction sites is a particularly challenging problem due to the presence o…

Pose EstimationPosition