paper-with-me

Papers

VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments

2026-02-23 · Jingyi Xu, Zhangshuo Qi, Zhongmiao Yan, Xuyu Gao, Qianyun Jiao, Songpengcheng Xia, Xieyuanli Chen, Ling Pei arxiv

In autonomous driving, robust place recognition is critical for global localization and loop closure detection. While inter-modality fusion of camera and LiDAR data in multimodal place recognition (MPR) has shown promise in overcoming the limitations of unimodal counterparts, existing MPR methods basically attend to hand-crafted fusion strategies and heavily parameterized backbones that require costly retraining. To address this, we propose VGGT-MPR, a multimodal place recognition framework that adopts the Visual Geometry Grounded Transformer (VGGT) as a unified geometric engine for both global retrieval and re-ranking. In the global retrieval stage, VGGT extracts geometrically-rich visual embeddings through prior depth-aware and point map supervision, and densifies sparse LiDAR point clouds with predicted depth maps to improve structural representation. This enhances the discriminative ability of fused multimodal features and produces global descriptors for fast retrieval. Beyond global retrieval, we design a training-free re-ranking mechanism that exploits VGGT's cross-view keypoint-tracking capability. By combining mask-guided keypoint extraction with confidence-aware correspondence scoring, our proposed re-ranking mechanism effectively refines retrieval results without additional parameter optimization. Extensive experiments on large-scale autonomous driving benchmarks and our self-collected data demonstrate that VGGT-MPR achieves state-of-the-art performance, exhibiting strong robustness to severe environmental changes, viewpoint shifts, and occlusions. Our code and data will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2602.19735

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingPoint Clouds

Similar Papers 제목 키워드 기반

VGGT-SLAM++

2026-04-08 · Avilasha Mandal, Rajesh Kumar, Sudarshan Sunil Harithas, Chetan Arora arxiv

We introduce VGGT-SLAM++, a complete visual SLAM system that leverages the geometry-rich outputs of the Visual Geometry Grounded Transformer (VGGT). The system comprises a visual odometry (front-end) fusing the VGGT feed…

Visual Place RecognitionVisual Odometry

SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes

2025-11-23 · Jungho Lee, Minhyeok Lee, Sunghun Yang, Minseok Kang 외 arxiv

3D reconstruction in large-scale scenes is a fundamental task in 3D perception, but the inherent trade-off between accuracy and computational efficiency remains a significant challenge. Existing methods either prioritize…

Visual Place RecognitionComputational Efficiency3D Reconstruction

OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

2025-11-13 · Haosong Peng, Hao Li, Yalun Dai, Yushi Lan 외 arxiv

General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). …

Camera Pose EstimationDepth Estimation

VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation

2026-03-19 · Jiayi Yuan, Haobo Jiang, De Wen Soh, Na Zhao arxiv

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panor…

3D ReconstructionDepth Estimation

VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation

2026-04-15 · Yulu Gao, Bohao Zhang, Zongheng Tang, Jitong Liao 외 arxiv

Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for applications in embodied AI and remote collaboration. This task is exce…

Semantic SegmentationObject Segmentation