paper-with-me

홈 › Papers

VISTA: Monocular Segmentation-Based Mapping for Appearance and View-Invariant Global Localization

2025-07-15 · Hannah Shafferman, Annika Thomas, Jouko Kinnari, Michael Ricard, Jose Nino, Jonathan How arxiv

Global localization is critical for autonomous navigation, particularly in scenarios where an agent must localize within a map generated in a different session or by another agent, as agents often have no prior knowledge about the correlation between reference frames. However, this task remains challenging in unstructured environments due to appearance changes induced by viewpoint variation, seasonal changes, spatial aliasing, and occlusions -- known failure modes for traditional place recognition methods. To address these challenges, we propose VISTA (View-Invariant Segmentation-Based Tracking for Frame Alignment), a novel open-set, monocular global localization framework that combines: 1) a front-end, object-based, segmentation and tracking pipeline, followed by 2) a submap correspondence search, which exploits geometric consistencies between environment maps to align vehicle reference frames. VISTA enables consistent localization across diverse camera viewpoints and seasonal changes, without requiring any domain-specific training or finetuning. We evaluate VISTA on seasonal and oblique-angle aerial datasets, achieving up to a 69% improvement in recall over baseline methods. Furthermore, we maintain a compact object-based map that is only 0.6% the size of the most memory-conservative baseline, making our approach capable of real-time implementation on resource-constrained platforms.

📄 PDF Abstract BibTeX arXiv:2507.11653

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaffold-SLAM: Structured 3D Gaussians for Simultaneous Localization and Photorealistic Mapping

2025-01-09 · Wen Tianci, Liu Zhiang, Lu Biao, Fang Yongchun

3D Gaussian Splatting (3DGS) has recently revolutionized novel view synthesis in the Simultaneous Localization and Mapping (SLAM). However, existing SLAM methods utilizing 3DGS have failed to provide high-quality novel v…

3DGSNovel View SynthesisSimultaneous Localization and Mapping

Vista4D: Video Reshooting with 4D Point Clouds

2026-04-23 · Kuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant 외 arxiv

We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the s…

Depth EstimationPoint Clouds

VISTA: A Panoramic View of Neural Representations

2024-12-03 · Tom White

We present VISTA (Visualization of Internal States and Their Associations), a novel pipeline for visually exploring and interpreting neural network representations. VISTA addresses the challenge of analyzing vast multidi…

ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association

2025-09-01 · Ganlin Zhang, Shenhan Qian, Xi Wang, Daniel Cremers arxiv

We present ViSTA-SLAM as a real-time monocular visual SLAM system that operates without requiring camera intrinsics, making it broadly applicable across diverse camera setups. At its core, the system employs a lightweigh…

3D Reconstruction

Semi-Dense 3D Semantic Mapping from Monocular SLAM

2016-11-13 · Xuanpeng Li, Rachid Belaroussi

The bundle of geometry and appearance in computer vision has proven to be a promising solution for robots across a wide variety of applications. Stereo cameras and RGB-D sensors are widely used to realise fast 3D reconst…

3D ReconstructionSemantic Segmentation