paper-with-me

Papers

M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM

2026-03-17 · Kerui Ren, Guanghao Li, Changjian Jiang, Yingxiang Xu, Tao Lu, Linning Xu, Junting Dong, Jiangmiao Pang, Mulin Yu, Bo Dai arxiv

Streaming reconstruction from uncalibrated monocular video remains challenging, as it requires both high-precision pose estimation and computationally efficient online refinement in dynamic environments. While coupling 3D foundation models with SLAM frameworks is a promising paradigm, a critical bottleneck persists: most multi-view foundation models estimate poses in a feed-forward manner, yielding pixel-level correspondences that lack the requisite precision for rigorous geometric optimization. To address this, we present M^3, which augments the Multi-view foundation model with a dedicated Matching head to facilitate fine-grained dense correspondences and integrates it into a robust Monocular Gaussian Splatting SLAM. M^3 further enhances tracking stability by incorporating dynamic area suppression and cross-inference intrinsic alignment. Extensive experiments on diverse indoor and outdoor benchmarks demonstrate state-of-the-art accuracy in both pose estimation and scene reconstruction. Notably, M^3 reduces ATE RMSE by 64.3% compared to VGGT-SLAM 2.0 and outperforms ARTDECO by 2.11 dB in PSNR on the ScanNet++ dataset.

📄 PDF Abstract BibTeX arXiv:2603.16844

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

Flow Matching Meets Biology and Life Science: A Survey

2025-07-23 · Zihao Li, Zhichen Zeng, Xiao Lin, Feihao Fang 외 arxiv

Over the past decade, advances in generative modeling, such as generative adversarial networks, masked autoencoders, and diffusion models, have significantly transformed biological research and discovery, enabling breakt…

Drug Discovery

VGGT-X: When VGGT Meets Dense Novel View Synthesis

2025-09-29 · Yang Liu, Chuanchen Luo, Zimo Tang, Junran Peng 외 arxiv

We study the problem of applying 3D Foundation Models (3DFMs) to dense Novel View Synthesis (NVS). Despite significant progress in Novel View Synthesis powered by NeRF and 3DGS, current approaches remain reliant on accur…

Novel View SynthesisPose EstimationPoint Clouds

Capture Dense: Markerless Motion Capture Meets Dense Pose Estimation

2018-12-05 · Xiu Li, Yebin Liu, Hanbyul Joo, Qionghai Dai 외

We present a method to combine markerless motion capture and dense pose feature estimation into a single framework. We demonstrate that dense pose information can help for multiview/single-view motion capture, and multiv…

Human ParsingMarkerless Motion CapturePose Estimation

When Epipolar Constraint Meets Non-local Operators in Multi-View Stereo

2023-09-29 · ICCV 2023 1 · Tianqi Liu, Xinyi Ye, Weiyue Zhao, Zhiyu Pan 외

Learning-based multi-view stereo (MVS) method heavily relies on feature matching, which requires distinctive and descriptive representations. An effective solution is to apply non-local feature aggregation, e.g., Transfo…

3D ReconstructionDescriptivePoint CloudsStereo Matching

Dense-SfM: Structure from Motion with Dense Consistent Matching

2025-01-24 · CVPR 2025 1 · Jongmin Lee, Sungjoo Yoo

We present Dense-SfM, a novel Structure from Motion (SfM) framework designed for dense and accurate 3D reconstruction from multi-view images. Sparse keypoint matching, which traditional SfM methods often rely on, limits …

3D Reconstruction