paper-with-me

Papers

VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction

2026-01-27 · Dominic Maggio, Luca Carlone arxiv

We present VGGT-SLAM 2.0, a real-time RGB feed-forward SLAM system which substantially improves upon VGGT-SLAM for incrementally aligning submaps created from VGGT. Firstly, we remove high-dimensional 15-degree-of-freedom drift and planar degeneracy from VGGT-SLAM by creating a new factor graph design while still addressing the reconstruction ambiguity of VGGT given unknown camera intrinsics. Secondly, by studying the attention layers of VGGT, we show that one of the layers is well suited to assist in image retrieval verification for free without additional training, which enables both rejecting false positive matches and allows for completing more loop closures. Finally, we conduct a suite of experiments which includes showing VGGT-SLAM 2.0 can easily be adapted for open-set object detection and demonstrating real-time performance while running online onboard a ground robot using a Jetson Thor. We test in environments ranging from cluttered indoor apartments and office scenes to a 4,200 square foot barn, and we also demonstrate VGGT-SLAM 2.0 achieves the highest accuracy on the TUM dataset with about 23 percent less pose error than VGGT-SLAM. Code will be released upon publication.

📄 PDF Abstract BibTeX arXiv:2601.19887

Code (0)

등록된 구현이 없습니다.

Tasks

Object DetectionImage Retrieval

Similar Papers 제목 키워드 기반

HyVGGT-VO: Tightly Coupled Hybrid Dense Visual Odometry with Feed-Forward Models

2026-04-02 · Junxiang Pan, Lipu Zhou, Baojie Chen arxiv

Dense visual odometry (VO), which provides pose estimation and dense 3D reconstruction, serves as the cornerstone for applications ranging from robotics to augmented reality. Recently, feed-forward models have demonstrat…

Computational Efficiency3D ReconstructionPose EstimationVisual Odometry

VGGT-SLAM++

2026-04-08 · Avilasha Mandal, Rajesh Kumar, Sudarshan Sunil Harithas, Chetan Arora arxiv

We introduce VGGT-SLAM++, a complete visual SLAM system that leverages the geometry-rich outputs of the Visual Geometry Grounded Transformer (VGGT). The system comprises a visual odometry (front-end) fusing the VGGT feed…

Visual Place RecognitionVisual Odometry

VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

2025-05-18 · Dominic Maggio, Hyungtae Lim, Luca Carlone

We present VGGT-SLAM, a dense RGB SLAM system constructed by incrementally and globally aligning submaps created from the feed-forward scene reconstruction approach VGGT using only uncalibrated monocular cameras. While r…

GPU

SceneVGGT: VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation

2026-02-12 · Anna Gelencsér-Horváth, Gergely Dinya, Dorka Boglárka Erős, Péter Halász 외 arxiv

We present SceneVGGT, a spatio-temporal 3D scene understanding framework that combines SLAM with semantic mapping for autonomous and assistive navigation. Built on VGGT, our method scales to long video streams via a slid…

Scene UnderstandingChange DetectionSemantic SLAM

M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM

2026-03-17 · Kerui Ren, Guanghao Li, Changjian Jiang, Yingxiang Xu 외 arxiv

Streaming reconstruction from uncalibrated monocular video remains challenging, as it requires both high-precision pose estimation and computationally efficient online refinement in dynamic environments. While coupling 3…

Pose Estimation