paper-with-me

홈 › Papers

Loc$^2$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching

2025-09-11 · Zimin Xia, Chenghao Xu, Alexandre Alahi arxiv

We propose an accurate and interpretable fine-grained cross-view localization method that estimates the 3 Degrees of Freedom (DoF) pose of a ground-level image by matching its local features with a reference aerial image. Unlike prior approaches that rely on global descriptors or bird's-eye-view (BEV) transformations, our method directly learns ground-aerial image-plane correspondences using weak supervision from camera poses. The matched ground points are lifted into BEV space with monocular depth predictions, and scale-aware Procrustes alignment is then applied to estimate camera rotation, translation, and optionally the scale between relative depth and the aerial metric space. This formulation is lightweight, end-to-end trainable, and requires no pixel-level annotations. Experiments show state-of-the-art accuracy in challenging scenarios such as cross-area testing and unknown orientation. Furthermore, our method offers strong interpretability: correspondence quality directly reflects localization accuracy and enables outlier rejection via RANSAC, while overlaying the re-scaled ground layout on the aerial image provides an intuitive visual cue of localization performance.

📄 PDF Abstract BibTeX arXiv:2509.09792

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Leveraging Previous-Traversal Point Cloud Map Priors for Camera-Based 3D Object Detection and Tracking

2026-04-28 · Markus Käppeler, Özgün Çiçek, Yakov Miron, Abhinav Valada arxiv

Camera-based 3D object detection and tracking are central to autonomous driving, yet precise 3D object localization remains fundamentally constrained by depth ambiguity when no expensive, depth-rich online LiDAR is avail…

Object Localization3D Object DetectionAutonomous Driving

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes

2026-06-30 · Kartik Bali, Roland Aydin arxiv

Identifying and grounding precise geometric entities, such as edges, planar regions, and curved surfaces within 3D objects, is foundational to computer-aided design (CAD), robotic manipulation, and scientific simulation.…

Natural Language Queries

UniSem: Generalizable Semantic 3D Reconstruction from Sparse Unposed Images

2026-03-18 · Guibiao Liao, Qian Ren, Kaimin Liao, Hua Wang 외 arxiv

Semantic-aware 3D reconstruction from sparse, unposed images remains challenging for feed-forward 3D Gaussian Splatting (3DGS). Existing methods often predict an over-complete set of Gaussian primitives under sparse-view…

3D ReconstructionDepth Estimation

Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks

2022-05-26 · Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, Yaoyu Zhang

Understanding the relation between deep and shallow neural networks is extremely important for the theoretical study of deep learning. In this work, we discover an embedding principle in depth that loss landscape of an N…

DCHM: Depth-Consistent Human Modeling for Multiview Detection

2025-07-19 · Jiahao Ma, Tianyu Wang, Miaomiao Liu, David Ahmedt-Aristizabal 외 arxiv

Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for …

Pedestrian DetectionMultiview DetectionDepth EstimationPoint Clouds