paper-with-me

Papers

VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

2025-05-18 · Dominic Maggio, Hyungtae Lim, Luca Carlone

We present VGGT-SLAM, a dense RGB SLAM system constructed by incrementally and globally aligning submaps created from the feed-forward scene reconstruction approach VGGT using only uncalibrated monocular cameras. While related works align submaps using similarity transforms (i.e., translation, rotation, and scale), we show that such approaches are inadequate in the case of uncalibrated cameras. In particular, we revisit the idea of reconstruction ambiguity, where given a set of uncalibrated cameras with no assumption on the camera motion or scene structure, the scene can only be reconstructed up to a 15-degrees-of-freedom projective transformation of the true geometry. This inspires us to recover a consistent scene reconstruction across submaps by optimizing over the SL(4) manifold, thus estimating 15-degrees-of-freedom homography transforms between sequential submaps while accounting for potential loop closure constraints. As verified by extensive experiments, we demonstrate that VGGT-SLAM achieves improved map quality using long video sequences that are infeasible for VGGT due to its high GPU requirements.

📄 PDF Abstract BibTeX arXiv:2505.12549

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

VGGT-SLAM++

2026-04-08 · Avilasha Mandal, Rajesh Kumar, Sudarshan Sunil Harithas, Chetan Arora arxiv

We introduce VGGT-SLAM++, a complete visual SLAM system that leverages the geometry-rich outputs of the Visual Geometry Grounded Transformer (VGGT). The system comprises a visual odometry (front-end) fusing the VGGT feed…

Visual Place RecognitionVisual Odometry

VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction

2026-01-27 · Dominic Maggio, Luca Carlone arxiv

We present VGGT-SLAM 2.0, a real-time RGB feed-forward SLAM system which substantially improves upon VGGT-SLAM for incrementally aligning submaps created from VGGT. Firstly, we remove high-dimensional 15-degree-of-freedo…

Object DetectionImage Retrieval

VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency

2026-02-05 · Zhuang Xiong, Chen Zhang, Qingshan Xu, Wenbing Tao arxiv

Despite recent progress in calibration-free monocular SLAM via 3D vision foundation models, scale drift remains severe on long sequences. Motion-agnostic partitioning breaks contextual coherence and causes zero-motion dr…

AIM-SLAM: Dense Monocular SLAM via Adaptive and Informative Multi-View Keyframe Prioritization with Foundation Model

2026-03-05 · Jinwoo Jeon, Dong-Uk Seo, Eungchang Mason Lee, Hyun Myung arxiv

Recent advances in geometric foundation models have emerged as a promising alternative for addressing the challenge of dense reconstruction in monocular visual simultaneous localization and mapping (SLAM). Although geome…

Pose Estimation

M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM

2026-03-17 · Kerui Ren, Guanghao Li, Changjian Jiang, Yingxiang Xu 외 arxiv

Streaming reconstruction from uncalibrated monocular video remains challenging, as it requires both high-precision pose estimation and computationally efficient online refinement in dynamic environments. While coupling 3…

Pose Estimation